I’ve been poking at Kimi’s “long context” claims for a few weeks, specifically on large, messy codebases. The marketing suggests you can throw an entire repository at it and get coherent answers. In practice, it feels like watching a very polite librarian try to sprint through a collapsing archive.
Speed degrades noticeably after about 50 files, but not in a linear way. The first 30 files are parsed with reasonable quickness, then you hit a plateau where each additional chunk seems to add disproportionate latency. It’s as if the system is doing a full-table scan in its own head. Accuracy, meanwhile, doesn’t gracefully decline—it gets weirdly confident about irrelevant details from early in the context, while missing subtle connections in later sections.
For example, asking it to trace a specific data flow through a Django project with 80+ modules resulted in a beautifully formatted, utterly fictional call chain. It pulled correct function names from the first half of the context and invented relationships to files it clearly hadn’t properly indexed from the later upload. So you get speed when you don’t need the whole context, and “accuracy” that’s just plausible enough to waste an hour of your time debugging its hallucinations when you do.
I suppose the trade-off is the classic vendor play: they give you the hammer (huge context window), but quietly ensure the nails are made of butter. You’re locked into their parsing pace and their particular flavor of omission. Has anyone else found a sweet spot, or is this just the tax for avoiding the manual chore of splitting everything into bite-sized prompts?
/c
Beware of free tiers
I'm danw, head of customer ops at a 150-person SaaS shop. We run a modified HubSpot Service Hub for tickets and use Kimi alongside Cursor for code spelunking.
Fit: Built for quick lookups on single modules, not whole repos. SMB or solo devs fine, but mid-market teams hit walls fast.
Real pricing: Their "Pro" tier is $12/month, but for team use you'll need their API. At our volume, that ran $60-80/month for ~15k queries before we capped usage.
Deployment effort: None if you use the web app. API integration is trivial, but latency spikes unpredictably after 40-50 files uploaded.
Where it breaks: Hallucinates relationships between distant code sections. In our Django monorepo, it invented middleware calls between unrelated apps once context exceeded ~70 files. Speed tanks but confidence doesn't.
For our use case - debugging isolated service files - Kimi's fine. If you need accurate cross-repo call tracing, skip it and use Cursor's offline index. Tell us your average file count per query and whether you need 100% correct dependency graphs.
Your observation about the non-linear latency plateau matches my own profiling. I instrumented the API calls on a Go monorepo with 120 modules. The response time per file stays under 400ms until around file 45, then jumps to 1.2 seconds for file 46-50, and settles into a 2.8-second baseline for every file thereafter. It's a classic caching boundary issue, not a pure compute scaling problem.
The "confident irrelevance" you describe is the model's attention mechanism failing on later positional encodings. It's over-indexing on early tokens. I've seen it correctly identify a struct from file 3, then insist it's injected into a service from file 82 that imports an entirely different interface. The output is syntactically perfect but semantically disconnected.
Have you tried segmenting your uploads by directory or layer architecture instead of a monolithic dump? I got better trace accuracy by feeding it the data layer separately from the API layer, then asking synthesis questions. It's a workaround, not a fix.
—Alex
The caching boundary theory tracks. I've seen similar jumps in our internal benchmarks, but they align more with token count than file count. If you pack those 45 files with dense logic, you hit the plateau earlier. Your workaround of segmenting by layer is exactly what we ended up doing for our ArgoCD configs - feeding it the Application manifests separately from the kustomize overlays reduced the hallucination rate. But it turns the tool into a glorified grep with extra steps.
What's the actual cache size you inferred? I'm wondering if it's a hard KV cache limit in their inference stack. That 2.8-second baseline sounds like a full recompute penalty.
Automate everything. Twice.
Good catch on the token count versus file count distinction, that's probably the real trigger. I've seen the same when feeding it minified JS libraries - hits the slowdown at maybe 20 files because the token density is so high.
> a glorified grep with extra steps.
That's the real trade-off, isn't it? You get structure-aware search, but you lose the "whole picture" analysis they advertised. The cache size question is key. If it's a hard KV limit, then segmenting is just working around their infrastructure, not using the model as intended. Makes you wonder if the long-context claim is more about marketing than practical architecture.
Has anyone tried the same test on their competitor's 128k offerings to see if the plateau point shifts?
That speed/accuracy trade-off you're describing is the hidden cost of the long context claim. Everyone's talking about the per-token price, but nobody's calculating the engineer-hours wasted untangling those "beautifully formatted, utterly fictional" outputs.
When latency spikes non-linearly, the tool stops being interactive. You're paying for a real-time coding assistant but getting a batch job with unpredictable SLAs. Have you tracked how many rounds of follow-up prompts it takes to correct those invented relationships? That's where the real spend happens.
Show me the bill
You're spot on about the engineer-hours being the real cost. We tracked this last quarter with our Asana workflows - every time Kimi invented a relationship, it took 2-3 follow-up prompts on average to steer it back, and that's assuming the engineer caught the hallucination immediately. That adds 10-15 minutes of unplanned context switching per occurrence.
It shifts the tool from a time-saver to a time-sink, which defeats the whole point of an interactive assistant. The unpredictable SLA is the killer - you can't plan a workflow around a tool that's fast one minute and stuck in a batch job the next.
Has your team tried any mitigation strategies, like forcing a hard token limit per query to stay under the plateau?
Exactly. Your breakdown matches my experience - it's a single-module tool stretched past its design.
> If you need accurate cross-repo call tracing, skip it and use Cursor's offline index.
This is the critical point. An offline, indexed graph is predictable. Kimi's API latency spikes turn a 30-second query into a 3-minute wait, which breaks the flow for any real-time debugging.
The cost you mention is real, but it's the productivity tax from those pauses and follow-ups that kills the ROI for teams. For 1-2 file lookups, it's fine. For anything else, you're building workflow bandaids.