Skip to content
Notifications
Clear all

Help: Anthropic's messages API timing out on 16k context fills.

3 Posts
3 Users
0 Reactions
31 Views
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
Topic starter   [#23416]

Hitting consistent timeouts with Anthropic's messages API when we approach ~16k context. Using claude-3-haiku-20240307. Happens reliably on tasks like summarizing large code diffs.

Our setup:
- Direct API calls, no streaming
- Simple system prompt + user content filling context
- Timeout set to 120s, but we're hitting it around 90s
- Smaller contexts (<8k) work fine

Seems like a sharp cliff. Anyone else hitting this? Specifically:
- Is this a model-specific issue with Haiku?
- Are there undocumented request size or processing time limits?
- Any workarounds besides chunking below 16k?

Looking for practical fixes, not "use a different provider." Need to make this reliable.



   
Quote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

We hit this on our CI summaries too. Same model, same timeout wall around 15-16k tokens. It's definitely not just you.

Switched to streaming and the problem vanished. The timeout seems tied to waiting for the full response buffer. With streaming, you get the first token fast and keep the connection alive. It's a bit more code but fixed our reliability. Still annoying though.

Have you checked if it's the total prompt size or just the output generation that's slow? We saw similar delays even with a tiny max_tokens when the input was large.



   
ReplyQuote
(@emma88)
Reputable Member
Joined: 3 months ago
Posts: 208
 

So streaming fixed it for you? We tried that and still saw timeouts, but only on our cheaper VPS nodes. On better hardware it helped.

Our theory is the timeout gets triggered if the whole response takes too long to generate plus download. If your network is slow, streaming doesn't save you. Did you compare response times between your environments?



   
ReplyQuote