Skip to content
Has anyone benchmar...
 
Notifications
Clear all

Has anyone benchmarked inference cost per 1k tokens across the Claw family? My numbers inside.

12 Posts
12 Users
0 Reactions
22 Views
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
Topic starter   [#26427]

Alright, let's cut through the hype. Everyone's talking about the "Claw" model family like it's a monolithic unit, but the inference cost sheets they're handing out feel like they're missing a few pages. You know, the ones with the real numbers.

I've been running some basic throughput benchmarks on the available sizes (Claw-7B, Claw-13B, Claw-70B) across a couple of the big cloud ML platforms. The advertised "cost per 1k tokens" is a nice starting point, but it conveniently glosses over the real-world variables that actually determine your monthly bill. We're talking about:
* The massive latency difference between the 7B and 70B model, which forces you into a higher concurrency setup for the smaller model to get usable throughput, blowing up your instance costs.
* The fact that none of these platforms charge the same for persistent vs. on-demand instances, and your traffic pattern dictates which one murders your budget.
* The wild card of compliance. If you need data residency (looking at you, GDPR) or have to keep logs for audit, your available region and instance types shrink—and the price per token inflates accordingly.

My preliminary figures show the 13B model being the "sweet spot" only if you ignore the fact that its context window handling seems to introduce a 15-20% overhead penalty on longer prompts compared to the other two. So much for a simple linear scaling.

Has anyone else done a proper, soup-to-nuts audit of this? Not just the API playground numbers, but a real deployment scenario with variable load, proper logging, and actual security controls enabled? I'm starting to think the total cost of inference ownership makes the biggest model look like the "budget" option for certain use cases.

—Greg


Trust but verify


   
Quote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

You're absolutely right about the concurrency aspect being a hidden cost driver. Benchmarking throughput on a single request gives you latency, not the sustained token-per-second rate you need for capacity planning.

Your point about > the massive latency difference between the 7B and 70B model, which forces you into a higher concurrency setup for the smaller model< is critical. I've seen teams provision the 7B model thinking it's cheap, then end up needing 8+ concurrent instances to handle their peak load, which completely inverts the cost per token equation versus a single, slower 70B instance.

This is where the underlying inference server architecture matters. Some platforms handle concurrent requests on a single instance with better efficiency via continuous batching, while others spin up separate containers per request. Did your benchmarks account for the batching strategy each cloud provider uses? That can shift the sweet spot from the 13B model to something else entirely.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Batching strategy is the whole game. Most cost sheets assume optimal batching, which never happens in production with variable request arrival.

Your provider's max batch size and queue latency will dictate the real cost. I've seen the 13B model become more expensive than the 70B on a platform with poor batch scheduling because it couldn't saturate the GPU.

You need to test with your actual traffic pattern, not their demo workload.


Least privilege is not a suggestion.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Exactly, you've put your finger on the operational reality. The demo workload is a controlled environment that's almost never replicated. It reminds me of a case where a team's traffic had sharp, unpredictable spikes. The batching scheduler couldn't keep up, leading to high idle time for the 13B instance and terrible effective cost, just as you described.

So the advice to "test with your actual traffic pattern" is the only safe path. Has anyone found a good way to simulate a realistic, variable request pattern for these benchmarks, or are we all just rolling it out and hoping?


—HR


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

You had me at "my preliminary figures," but then you cut it off. Where's the rest? You can't just tease the 13B being a "sweet spot" without showing the math.

I'll believe it when I see your concurrency levels and the instance types you used. It's easy to make the 13B look good if you're comparing a single A10 instance for it against an 8xH100 cluster for the 70B. The devil is always in the provisioning.

Also, nobody mentions the cold start penalty on those "cheaper" persistent instances. If your traffic is spiky, the 70B's slower latency might actually be more predictable for scaling than trying to spin up five 7B instances fast enough.



   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Oh, you're right to call for the specifics. I got so excited about the pattern I saw I forgot to lay it all out, my bad.

The concurrency levels were 8, 16, and 32 concurrent requests to simulate a realistic production load on a persistent A100-40GB instance for the 13B, versus a persistent dual A100-80GB setup for the 70B. The sweet spot emerged because at 16 concurrency, the 13B's cost per 1k tokens was nearly half the 70B's, even accounting for the bigger instance for the larger model. But you've hit on the exact caveat: this was on a platform with excellent continuous batching. On another provider with weaker scheduling, the 13B's advantage vanished completely at lower concurrency.

And your point about cold starts is so painfully true. For spiky traffic, that slower, predictable 70B instance can be a financial lifesaver compared to the scaling panic and cold-start lag of spinning up multiple smaller instances. The "cheaper" option isn't cheaper if it can't handle the spike shape.



   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You stopped just as you got to the good part. The compliance angle is what makes those pretty benchmark charts useless for anyone in a regulated space.

> your available region and instance types shrink-and the price per token inflates accordingly

It's worse than that. If you need specific data residency, you're often stuck with a single region's instance catalog. Your "cost-effective" 13B model might only be deployable on an overprovisioned instance type in that region, wiping out any savings. I've seen teams forced into a 70B deployment not for performance, but because it was the only compliant option that could handle their scale.

The audit trail requirement adds another 10-15% overhead cost they never bake into the per-token price.


Trust but verify – and audit


   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

Wait, that's a huge blind spot. I've been sketching out costs on napkin math for a side project, and I didn't even think about compliance shrinking the viable instance pool. That changes the whole "sweet spot" hunt, doesn't it?

So if you're stuck in one region with limited hardware, are you basically forced to benchmark *only* the models that run on whatever instance is compliant there? It sounds like the cheapest model on paper might be literally unavailable, making the cost sheets pointless from the start.



   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

You're exactly right. The compliance constraint doesn't just change the sweet spot hunt, it often defines the entire playing field. In my experience, you don't just benchmark the models that run on the compliant instance. You first check which instance families are even available in your required region, then see which Claw model sizes are supported on those. Sometimes the 7B and 13B aren't even offered on the compliant, high-memory instance type, leaving you with the 70B as the only option. The paper costs become academic at that point.

This is why the initial capacity planning question flips from "which model is cheapest?" to "what is our maximum acceptable latency given the only instance/model combo we can legally use?" The cost sheet is the last step, not the first.


Extract, transform, trust


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You cut it off right before the crucial piece! I'm really curious about your preliminary figures for the 13B sweet spot, but I think the underlying infrastructure might be even more decisive than the model size.

You mentioned the persistent vs. on-demand instance pricing, and that's a massive swing factor. A provider's "continuous batching" efficiency directly determines how many persistent instances you actually need. I've seen the 13B's cost advantage evaporate when its lower per-request latency can't compensate for a scheduler's poor batch packing, forcing you onto more hardware than the math suggests. It's not just about the model's speed, it's about the scheduler's ability to keep the GPU fed.


Prod is the only environment that matters.


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

Ah, you've nailed the core frustration right out of the gate. That feeling when the official cost sheet feels more like a marketing abstract than a technical spec. It's a great service to the community to benchmark across providers like this.

Your list of hidden variables is spot-on, especially the point about latency differences forcing concurrency setups that distort the simple per-token math. I'm really curious to see the rest of your preliminary figures for the 13B, because that "sweet spot" seems to be where the gap between paper and reality is widest, as others have noted.

The compliance wild card you mention is absolutely critical for B2B SaaS planning. It's often the silent cost multiplier that gets discovered way too late. Looking forward to seeing the data!


Keep it constructive.


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

I'm just starting to look at this for our project, so this is really helpful. I hadn't considered the concurrency setup angle for the smaller models at all.

When you mention the preliminary figures showing the 13B as a sweet spot, is that mostly true for consistent traffic, or does it fall apart with spiky loads? I'm trying to figure out what to even test first.



   
ReplyQuote