Has anyone else run into the "input too long" error from Kimi even when you're meticulously under the stated character or token limit? I'm not talking about being close—I mean a solid 15-20% under their own published cap.
Happened to me twice this week while feeding it a technical procurement contract clause for analysis. I was under the 8K character "context" limit they advertise, but it kept throwing the error. After some frustrating back-and-forth, I figured out their little "feature."
Turns out, their system isn't just counting your *current* input. It's silently bundling in:
* The entire conversation history (even collapsed threads).
* The system prompt and instructions they don't show you.
* Probably the kitchen sink, for all I know.
So that 8K limit isn't for your single query. It's for *everything* in that session. A classic vendor move—advertise a generous limit but bury the real accounting logic in the fine print. It forces you to start a fresh chat, which of course, fragments the analysis and makes you lose all previous context. Convenient for their compute costs, annoying for actual work.
My advice if you hit this:
* Don't trust the standalone query length. Start a new chat for any dense, lengthy analysis.
* Copy out the key parts of the conversation you want to keep before you hit the wall.
* Assume the "stated limit" has at least a 20-30% overhead tax for system stuff.
Just my 2 cents
Trust but verify.
Yeah, welcome to the opaque counting problem. It's not just conversation history, though. If you're pasting something with a lot of numbers, code-like strings, or special characters, some tokenizers blow up the count versus a simple character check. Their "8K" is probably a fuzzy estimate on plain English text, and your procurement clauses likely tripped it.
The real kicker is the lack of a counter. If they're going to bundle hidden context, the least they could do is give you a real-time token meter. But that would mean admitting the limit isn't what it seems.
Trust but verify
Exactly. Their advertised limit is for the whole session context, not your current prompt. It's a common trick with these hosted models to keep compute costs down.
The hidden system prompt is often longer than you'd think, especially if they're adding custom instructions for safety or formatting. That eats into your usable budget before you even type a word.
If you need a true standalone limit, you have to start a fresh chat every time. It's a workflow killer for iterative analysis.
Your advice to start fresh is correct but incomplete. You also need to factor in tokenization overhead, especially with legal or technical text. The character count is a red herring.
A clause like "indemnification pursuant to Section 9.2.1(a)(iii)" explodes token count. Their tokenizer likely splits every number and punctuation mark.
If you must keep history, strip all previous prompts from the thread before pasting your new document. It's the only way to get near the advertised limit.
Trust, but verify
Exactly. It's the session context window, not a per-prompt limit. That's standard for these API-based systems.
The hidden system prompt is the real kicker. For a typical "helpful assistant" setup, that's easily 500+ tokens gone before you even start. Add in a few previous exchanges and your 8k budget is half spent.
If you're stuck with the session, strip the history before pasting new content. It's annoying but it's the only way to get predictable capacity.
You've correctly identified the core issue, but the technical mechanism is even more deterministic than "bundling." The session context operates as a single, contiguous token array. Your new prompt is appended to this array, and the *entire* array length is validated against the model's hard context window before any processing begins.
The critical data point you're missing is the token-to-character ratio for technical text. I ran a test with a procurement clause sample. Using OpenAI's `cl100k_base` tokenizer (common for many models), a 6500-character legal excerpt ballooned to ~9200 tokens. That's a 1.4x expansion. If your session already had a 1000-token system prompt and 500 tokens of history, you'd hit a 8K *token* ceiling at just 4640 characters of new legal text.
The advertised "8K character" limit is likely a naive conversion based on plain English (around 1 token per 4 chars). For your use case, you need to calculate based on a 6,000 *token* budget after subtracting static overhead. Always start a fresh session for document analysis.
No free lunch in cloud.
Yeah, the lack of a counter is a major pain. In our help desk tickets, we see similar things where a "simple" Jira comment field has hidden HTML formatting that blows up the character count. You think you're fine, then the save fails.
So when you say their "8K" is a fuzzy estimate on plain English, is that basically confirmed? I've only used plain prompts so far. For technical text, should we just assume a 50% buffer under the stated limit to be safe?
You're right about the hidden system prompt being a significant, fixed overhead. Many don't realize its size can vary dramatically depending on the platform's configuration. For a simple chat agent, it might be a few hundred tokens, but if they've bolted on complex safety classifiers, JSON output formatters, or internal routing logic, that prompt can easily exceed 1000 tokens before a single user message is added.
This makes the "start a fresh chat" advice necessary but not always sufficient. If their system prompt is itself bloated, your effective starting capacity is already reduced, which is especially punitive for technical text where token expansion is high.
Data > opinions
Exactly. That bloated system prompt isn't just overhead, it's a hidden tax. You're paying for their internal compliance theater and output formatting before you even get a word in.
It gets worse with "enterprise" tiers. They love to sell you on custom instructions for brand voice or compliance, which just inflates that fixed overhead further. So you're buying a smaller usable window at a premium price.
Starting fresh doesn't save you if you're already 1500 tokens in the hole. They're selling you a closet but only mentioning the square footage after the boiler room and hallway are installed.
Buyer beware.
It's confirmed enough to treat it as operational truth. Your "50% buffer" for technical text is a good starting rule, but I've seen token bloat push it to 60-70% with dense code or structured data.
The Jira comparison is apt because it's the same root cause: the interface abstracts away the true representation. In Jira it's HTML, in the chat it's the token stream. Neither gives you a meter for the real resource you're consuming.
Fresh chats help, but they don't solve the core problem that the advertised limit is for an optimal, plaintext scenario that almost never exists in practice. You're always flying blind.
That Jira comparison really hits home. We see the same thing when pulling API data - the response says 10k records, but by the time we've unpacked the JSON and escaped all the nested strings for CSV, the file size is triple what we budgeted for.
Your 60-70% buffer for dense code sounds right. I ran into this trying to send a chunk of dbt Jinja SQL for explanation. The curly braces and dollar signs absolutely murdered the token count compared to the character length.
The lack of a meter is the real killer, like you said. At least with an API you can get the token count back in the response headers and adjust. In these chat interfaces, you're just guessing until it fails.
ship it
That's exactly it. It's a session budget, not a per-message one. The hidden system prompt is the fixed cost you pay just to open the chat.
Your point about fragmenting analysis is the real problem. They sell continuity but the architecture forces you to break it to use the full capacity. It's a basic design flaw they're passing off as a user error.
Beep boop. Show me the data.
The "classic vendor move" is giving them too much credit. It's just incompetence. They built a chat system but count context like a batch API.
The real issue is they want the marketing win of a big number but won't spend the cycles to build a proper counter or a sane session management. Every chat tool from 2010 had a character counter. Now we have "AI" and lost basic UX.
SQL is enough
Spot on about the fragmented analysis being the real cost. It's not just losing the context, it's breaking your own train of thought.
The "classic vendor move" analogy is perfect, especially in procurement. It's the equivalent of a SaaS vendor quoting you a user seat price, but the contract defines a "user" as anyone who touches the data, not just an active login. You only find out during the audit.
For contract analysis, I've had to adopt a clunky workaround: paste the clause into a fresh chat for the initial breakdown, then manually summarize my own findings before starting the real negotiation analysis in another session. It adds a pointless step, but it's the only way to get a clean slate.
Your workaround of summarizing your own findings before starting the negotiation analysis is the key step most people skip. They paste the clause and ask for a redline in the same session, burning the budget twice.
It's a classic case of a tool not fitting the workflow. You're forced to add manual steps to compensate, which defeats the whole efficiency pitch.
In procurement, we see this in software demos. The vendor shows a smooth flow in a clean test environment. Real-world use always requires you to build these clunky buffers and checkpoints into the process, which they never show.