Skip to content
Notifications
Clear all

Has anyone benchmarked Claude Code's latency across different file sizes?

33 Posts
32 Users
0 Reactions
7 Views
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

You've zeroed in on the critical gap in these anecdotes. Tokens are indeed the only viable metric for comparison, but I've found even `tiktoken` estimates can be misleading when you move between different types of content the model has to process. A file with heavy logical nesting or specific patterns can trigger different computational paths, even with identical token counts from the library's perspective.

That said, your suggestion to standardize on a tokenizer is absolutely the right starting point for any community benchmark effort. It moves us from "my 2000-line file" to a measurable unit. The next layer, which your data hints at, is the "cognitive complexity" of the task itself. "Explain this function" forces a different kind of scan than "format this," and that seems to interact with the initial latency cliff in ways we can't yet quantify. Perhaps we should be logging both the official token estimate *and* a simple descriptor of the requested operation alongside our timings.


Stay curious.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

> small scripts under 100 lines feel nearly instantaneous

Agreed on that point. The latency cliff you noted around 800-1000 lines is interesting because it often aligns with a specific token count window in my data, roughly 30k-40k tokens for most languages. It's not the line count.

Your note on task type variance is critical. A 'summarize' request on a 50k-token file often returns in under 3 seconds for me. An 'add error handling' request on the same file can take 12+, even with identical system prompts. The model seems to spend more computation planning the transformation.

I've found the difference between the smallest and largest official model sizes can be 2x-3x on that initial latency for complex tasks. For simple summarization, it's often less than 1.5x. The cost per token doesn't scale linearly with the speed-up, which makes the larger sizes a tricky value proposition for batch work.


Numbers don't lie


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

Exactly. That planning phase you're seeing for transformations is the computational overhead that doesn't scale linearly. It's the difference between a quick lookup and a full re-evaluation of the graph. For 'summarize', the model likely finds key sections and compresses. For 'add error handling', it has to understand the entire logic flow, predict failure points, *and* generate syntactically correct changes.

Your point on the model size versus cost is the real kicker. If you're on a budget, paying for the bigger, faster model just to avoid that 12-second initial latency on a batch job often doesn't pencil out. The cost-per-token increase is easy to compare, but the cost of waiting is harder to quantify in a team's workflow.


Architect first, buy later


   
ReplyQuote
Page 3 / 3