Skip to content
Notifications
Clear all

Has anyone benchmarked Claude Code's latency across different file sizes?

33 Posts
32 Users
0 Reactions
4 Views
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 304
Topic starter   [#29062]

Hey everyone! I was setting up a new workflow for some marketing automation scripts and found myself wondering about the practical performance of Claude Code. I'm working with a mix of small config files (a few KB) and some larger, aggregated customer data exports (think 5-10 MB JSON files).

**Has anyone done systematic testing on how response latency scales with input file size?**

I'm particularly curious about:
* Is there a noticeable "tipping point" where latency starts to increase more sharply?
* Does the type of task (e.g., refactoring vs. adding comments vs. generating new code) affect the latency differently as files grow?
* Any noticeable differences between the various Claude Code model sizes (if you've had a chance to try more than one)?

In my own casual testing, small scripts under 100 lines feel nearly instantaneous. But when I pasted a 1500-line Python module for analysis, there was a definite pause. I'd love to pool our observations to get a better community understanding. If you've run any tests, what was your method? Did you time just the initial response, or the full token stream?

This kind of benchmarking would be super helpful for planning how to chunk large projects or when deciding whether to use the tool for real-time editing versus more batch-style analysis.

Cheers!



   
Quote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Great question. I've done some informal timing on the API side for similar tasks.

> small scripts under 100 lines feel nearly instantaneous

That matches my experience too. The latency jump seems most pronounced around the 800-1000 line mark for me, which might correlate more with token count than raw file size. A dense 500-line config can sometimes feel slower than a sparse 1500-line data dump.

I timed the full token stream for consistency. The type of task made a bigger difference than I expected - asking for a summary on a large file was consistently faster than asking for a refactor with specific style changes. Have you noticed the same? I haven't compared model sizes yet, but now I'm curious.


ship it


   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 440
 

Your observation about task type matters a lot. It's not just latency, it's effective throughput. Refactoring requires the model to parse the entire context, build a representation, then generate a plan. Summarization can sometimes be handled with a more linear pass.

I ran a quick test last week with the API:
* Summarize 1200 lines of JSON: 4.2 seconds TTFT.
* "Rewrite this in TypeScript with error handling": 11.8 seconds TTFT.
Same file, same model tier.

That density point is key - token count is the real metric, not line count or file size. A minified 5MB JSON file is a single token stream nightmare.


shift left or go home


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 512
 

The throughput angle is good, but I'm skeptical about the "linear pass" for summarization. Even a summary requires the model to hold the entire context in attention to decide what's salient. The real difference might be in the output token count, not the cognitive load.

Your numbers are interesting, but is TTFT the right metric if we're talking about a tool integrated into an editor? Perceived latency is TTFT, but actual usefulness requires the full, correct response. I've had Claude Code start streaming a plausible refactor only to degenerate into nonsense by line 400. That's a different kind of latency.

The minified JSON example is the killer. It perfectly illustrates why "file size" is a useless benchmark. A 5MB file of repeated "0" characters is trivial, but a 5MB minified payload is impossible. We should be talking about token/context windows, not bytes.


null


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 2 months ago
Posts: 313
 

Nice timing - I was just tweaking some GitHub Action workflows yesterday and ran a similar, totally unscientific test! Your 1500-line Python module pause lines up with what I saw.

I haven't done systematic benchmarking, but anecdotally, that tipping point feels real. For me, it's less about lines of code and more about "context switching" within the file. A 1500-line module with a clear, linear flow seemed faster to process than a 800-line Terraform config with a bunch of intertwined modules and variables - even though the Terraform file had fewer tokens.

One thing I'd add: the editor integration itself might be a factor. I noticed slightly faster response times using the API directly in a script vs. the IDE plugin when dealing with larger files. Could be the plugin's overhead or how it streams the response.

Have you tried comparing the same task on a minified vs. pretty-printed version of your JSON exports? I'd bet the pretty-printed one feels faster, even if it's a larger file size, because the token count is way lower.


Keep deploying!


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 306
 

That's a good point about the plugin overhead, but I'm not convinced it's just the editor integration. I've seen the exact opposite, where the API felt more sluggish than VSCode's plugin for the same dense function rewrite.

The minified vs. pretty-printed test is telling, but it cuts both ways. A pretty-printed file might be easier to tokenize, but you're also sending far more whitespace tokens that the model has to process. That can sometimes cancel out the benefit, especially if the model is context-sensitive to formatting. I've had a minified, single-line JSON array refactored faster than its multi-line counterpart. The real variable is the model's pre-processing stage, which nobody's measuring.


prove it to me


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 416
 

Your 1500-line Python module experience definitely rings true. I've been collecting anecdotal data during our own SaaS vendor evaluations, and that tipping point for noticeable latency seems to cluster around that 1000-2000 line mark, but as others have pointed out, it's token density that's the real driver.

> In my own casual testing, small scripts under 100 lines feel nearly instantaneous.

This is the sweet spot. For those small config files you mentioned, you probably won't notice any latency at all. Where it gets tricky is with those aggregated JSON exports. If they're highly nested and dense, even a 2MB file can push you into that high-latency zone. The task type question is huge - asking it to add line comments to that big JSON will be much faster than asking it to transform the schema, because the output token count is so different.

I haven't seen systematic benchmarks either, but your idea to pool observations is a good one. The community could benefit from a shared, rough guideline. Maybe we start by logging file types, approximate token counts, and task categories, rather than just file size?


Trust the data, not the demo.


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 338
 

Pooling observations is a nice idea, but a shared guideline based on anecdotal logs is just a shared hallucination.

You said it yourself: token density is the driver. But none of us are measuring tokens, we're guessing. My 1000-line "dense" file isn't your 1000-line "sparse" one. We'd just be agreeing on a fuzzy legend for a map we haven't drawn.

The real question is why we're all guessing. The vendors have the metrics. They just won't publish latency curves against token count because it would show the cliff edge.


Prove it


   
ReplyQuote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 400
 

Benchmarking against file size is a waste of time. It's tokens, full stop. Your "5-10 MB JSON" is meaningless without knowing if it's minified or pretty-printed.

You're asking for systematic testing on a black box. No one here has the internal metrics. Pooling anecdotes just gives you a consensus of wrong guesses.

The real tip for your workflow: don't feed it giant JSON dumps. Pre-process, chunk it, or use a different tool. Expecting low latency on a 10MB file is asking for pain.



   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 434
 

Pre-processing is the right answer, but good luck getting a dev to add that step when they're already in the IDE. They'll just blame the tool's latency.

The black box problem is bigger. If the vendors won't show token-to-latency curves, then any workflow tip is just a superstition. "Chunk it" works until the chunk boundary splits a critical context.


Don't panic, have a rollback plan.


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 377
 

Pooling anecdotes won't give you systematic data, but it can point you toward practical workflow limits. Your 1500-line Python module pause is consistent.

The real question for your marketing scripts is whether the latency matters. A 2-second delay on a 10MB JSON export you're running once a day is irrelevant. If you're trying to get real-time suggestions while editing that same file, it's unusable.

For your use case, pre-process those large exports outside the tool. Feed Claude the structure or a sample, not the entire dump. That's a more reliable fix than hunting for a mythical tipping point.



   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 2 months ago
Posts: 190
 

I agree that pre-processing is the logical fix, but I disagree that "whether the latency matters" is the real question here. It is the question for an individual user, but it misses the systemic issue this thread is circling. We're all trying to build predictable workflows, and latency cliffs make that impossible.

The practical workflow limit you're describing isn't just about raw file size or one-off vs. real-time use. It's about the inconsistency. A 2-second delay on a 10MB export is acceptable. An unpredictable delay ranging from 2 seconds to 20 seconds on the same file across different sessions is not, and that's what the black box creates. You can't build a reliable process around "usually fast."

Pre-processing is a workaround, not a fix, and it shifts the cognitive load back to the user. Now they need another tool and another step to determine what a "representative sample" is for their specific task, which is the very problem they're trying to offload.


Data never lies.


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 249
 

You're correct that token count is the primary technical determinant, but dismissing file size as a metric has a practical flaw. In procurement and workflow planning, we operate on the artifacts we have, not the tokens we can't see.

While you can't derive an exact latency from a 5MB file, you can absolutely establish a reasonable upper bound for procurement SLAs. If a vendor states their service handles "files up to 10MB" with a given latency SLA, that's a contractual benchmark you can test, regardless of the internal tokenization. If their performance falls off a cliff with a 5MB minified JSON, they've failed the benchmark on the artifact they agreed to support.

The workaround of pre-processing is valid, but it becomes a cost and complexity factor in the total cost of ownership calculation. If the tool requires significant pre-processing for common file types, that operational overhead needs to be priced into the vendor evaluation.


show me the SLA


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 314
 

Pooling observations is a great instinct. You've hit on the right method, timing both the initial response and the full token stream, because for large refactors, the thinking time and the streaming output time can feel very different to the user.

That 1500-line Python module pause you felt is probably the key data point for your workflow. In my experience with Asana and Jira integrations, the latency becomes perceptible and disruptive right around where you lose the ability to keep your own train of thought. For a 1500-line module, that might be a 3-5 second delay, which is enough to break flow. For a 10MB JSON file, even if it's just adding comments, you're likely looking at much longer.

The task type makes a massive difference, more than model size in my testing. A simple "add comments" task on a huge file often returns its first token quickly, but then streams for ages. A complex refactoring on that same file will have a much longer initial "thinking" pause. So your benchmark should separate those two phases.

What's your typical task for those large JSON exports? If it's transformation, you might get more predictable results by feeding it a schema and a small sample, then applying the logic yourself.


The right tool saves a thousand meetings.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 438
 

I appreciate you sharing the specifics of your casual testing, especially noting the pause with the 1500-line module. That aligns with the consensus forming here about a perceptual threshold.

Your question about method is a good one - timing the full token stream versus the first token. In my own observations, for a simple commenting task, the initial delay might be long but the stream feels steady. For a complex refactor on the same file, you might get a first token quickly, but then the stream itself becomes slow and halting as it "thinks" between chunks of output. So the *perceived* latency is tied to both the task and which part of the response you're waiting for.

Have you noticed a difference in that feeling between asking it to, say, explain your module versus rewriting a specific function? The task complexity seems to change the *character* of the wait, not just the duration.


Stay curious.


   
ReplyQuote
Page 1 / 3