Skip to content
Notifications
Clear all

Practical question: How many RPS can I realistically expect from Claude Sonnet?

1 Posts
1 Users
0 Reactions
26 Views
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
Topic starter   [#13776]

Everyone's talking about tokens per second, but let's be honest: that's a vendor-friendly abstraction. What actually matters when you're trying to build something that works is requests per second. Specifically, sustained RPS without getting throttled into the next fiscal quarter.

So, the million-dollar question: what's the realistic, non-marketing RPS ceiling for Claude 3.5 Sonnet via their API? I'm not interested in the burst rate for the first three seconds. I mean a steady-state load over several minutes, the kind that would simulate actual user traffic.

I've seen the docs. They mention "tiered rate limits" and scaling with usage. That's a polite way of saying "it depends, and we won't tell you the number until you're already in production and hitting a wall." Has anyone actually pushed it? What did you get? 10 RPS? 50? Did you get a polite email from their trust and safety team disguised as a "capacity review"?

Bonus points for context: were you using a standard completion call or the messages API? What was your average token count per request? And most importantly, did the latency at the 95th percentile stay sane, or did it fall off a cliff once you passed a certain threshold?

I'm drafting a capacity plan and "trust us, it scales" isn't a unit of measurement.

cg


cg


   
Quote