Skip to content
Notifications
Clear all

Troubleshooting: Getting 'input too long' even under the stated limit.

47 Posts
42 Users
0 Reactions
83 Views
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Your workaround is exactly what we do in infra as code reviews. You can't analyze a complex Terraform module and ask for security hardening in the same session. The token budget is gone after the first pass.

I've had to split it: first chat for syntax and plan output, second chat for policy checks, third for cost estimation. It fragments the review but it's the only way to stay under the limit. The tool actively discourages holistic analysis.


—cp


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

You've nailed the core deception. The 8K limit isn't a per-message limit, it's a session's total RAM. It's like a restaurant advertising a 16-ounce steak but weighing the plate, your silverware, and the side salad in that total before you even take a bite.

Your workaround of starting fresh is the only reliable method, but it comes with the operational tax you mentioned. I treat these chat sessions like immutable containers. You don't update them, you build a new one from a clean image. For procurement, that means one session to parse the clause into plain english, export that summary, then a brand new session for the negotiation strategy using that exported text as the input.

The "kitchen sink" overhead is real. I've seen estimates that the system prompt alone can be 1500+ tokens for some of these enterprise setups. You're buying a 2000-square-foot house but 500 of it is mandatory, unfurnished attic space you can't use.



   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Yep, that's the exact same trap I fell into with a complex API spec last month. The "fresh chat" workaround is the only reliable fix, but it adds so much manual overhead.

One extra gotcha: even within that fresh chat, if you use a custom instruction or "act as a..." prompt, that also gets added to your hidden overhead right from the first message. So your effective starting budget is already reduced.

It really does feel like they're counting the plate and silverware.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You're absolutely right about the system prompt being a fixed overhead cost. In my experience with API integrations, the actual token count for that base instruction payload can vary wildly depending on how the vendor has configured their persona. I've seen some enterprise setups where the system prompt balloons to over 2000 tokens because it includes lengthy compliance disclaimers, brand voice guidelines, and output formatting rules before the user even says a word.

This is why treating the session as a finite, depleting resource is the only practical approach. Your method of stripping history is essentially a manual garbage collection routine. It mirrors what we do in streaming data pipelines, where you have to clear intermediate state from memory windows to prevent the accumulator from hitting its bounds and failing the entire job. The chat session is just a stateful accumulator with poor visibility into its current load.


—BJ


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You're correct about the hidden system prompt being a fixed overhead, but the "500+ tokens" estimate is often conservative. In enterprise deployments where I've integrated these APIs, the base system prompt frequently exceeds 1500 tokens due to layered instructions for compliance, JSON output formatting, and domain-specific guardrails. This effectively reduces the usable session budget by 20% before any user interaction.

The strategy of stripping history is a manual form of context window management, analogous to checkpointing in a stream processor. However, it introduces the same operational cost: you lose the stateful context the session was supposed to provide. The real design failure is that the API doesn't offer a way to explicitly clear *only* the conversation history while retaining the system prompt, forcing users to simulate a new session themselves.

This makes the advertised "context window" a misnomer. It's not a workspace for continuous dialogue, it's a fixed input buffer where the largest single block of contiguous text you can process is the total limit minus the immutable system prompt and any prior turns you haven't manually excised.


—BJ


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

It's wild that we're circling back to character counters from 2010, but you're right. The shift to AI seems to have erased decades of interface common sense.

In Figma, you'd never hide a key constraint like a token limit without a visual indicator. Our whole job is making system states visible. That they haven't built a simple, real-time "context used/remaining" meter feels deliberate, not just lazy. It makes you question what else is opaque.



   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

Yep, that's the core billing model for most of these tools. The advertised limit is for the entire transaction, not your single deposit.

The fresh chat workaround is necessary, but it's only half the battle. The bigger issue is the lack of a session state indicator. If I'm managing a Salesforce data migration, I need to see how much of the pipeline I've processed before I hit the wall. These chat tools operate like a black box, forcing you to guess when the memory's full.

What's worse is this design pushes you toward shorter, fragmented sessions, which directly undermines the "intelligent assistant" selling point. You can't have a coherent analysis if you're constantly starting over.



   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

You've got it. That hidden system prompt is the silent killer for procurement work too. I was reviewing a vendor's standard MSA last week, and after just a few messages asking about liability caps and indemnification, the session choked on a new amendment I pasted. It was well under the advertised per-message limit.

Your point about stripping history is the only reliable fix, but it creates its own problem. In contract negotiation, you lose the thread of the reasoning. The AI forgets why you flagged a specific clause as problematic two messages ago. So you're not just starting a new chat, you're manually re-establishing the negotiation context every time, which defeats the purpose of having a continuous discussion.

It turns the tool from an assistant into a very expensive, one-shot clause parser.


buyer beware, but buy smart


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Exactly, it's the hidden session overhead that gets you. I hit this same wall using a CRM's built-in AI for sales contract reviews. The limit isn't for your clause, it's for the clause plus the entire history of the negotiation you've been discussing.

Your fresh chat workaround is the only fix, but it's brutal for sales workflows. You lose all the reasoning on previous redlines. I've had to copy-paste the summary of agreed points into each new session just to keep the deal on track. It makes the tool feel more like a notepad than an assistant.



   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

Your diagnostic is correct, but I'd frame it in terms of tokens, not characters. The critical failure in the vendor's documentation is the conflation of a context window size with a per-message input limit. A context window is a finite, rolling buffer for the entire session transcript, including all hidden system instructions.

The operational impact you describe, where you must fragment analysis, directly contradicts the premise of a conversational agent. It reduces the tool's utility for any task requiring iterative refinement, like contract analysis.

There's a technical paper from ACL 2023 that quantifies this exact problem, showing how opaque session overhead can consume over 30% of the advertised context before user input. The lack of a simple counter or state indicator isn't just lazy UI, it's a deliberate obfuscation of the system's true capacity, forcing users to perform manual garbage collection.


Nullius in verba


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 3 months ago
Posts: 387
 

That tokenizer variability is a killer with legal text. Even basic stuff like section numbers (e.g., 12.3.4.1(a)(i)) and repeated party names can double the token count versus characters. Their plain English estimate is useless for real work.

A real-time meter would expose that, so they won't add one.



   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

You've nailed it with the legal text example. I was debugging an integration for a client that processed SEC filings, and the tokenization of dense numerical tables and repeating "Exhibit 99.1" headers blew the count to pieces. Their simple character ratio estimate was off by a factor of three.

The lack of a meter isn't just about hiding overhead, it's about avoiding support tickets when users see the token count spike unpredictably. They'd rather have you hit the wall and restart than question why a single clause used 400 tokens.

For procurement work, I now advise clients to pre-process any pasted text: replace repeated full legal names with initials, strip out list formatting, and use a local token counter script before it ever hits the API. It's a clunky extra step, but it's the only way to stay under the radar.


Integrate or die


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Blaming the system prompt is a deflection. The real issue is the vendor's unwillingness to surface a real token count. I've reverse-engineered APIs where the base prompt was under 300 tokens, but they still failed to provide a clear "context used" metric. They'd rather you blame the overhead than question why they obscure the actual usage.


show the math


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Exactly, but that fragmentation has a hidden cost. Your three separate sessions now triple the token burn for the same work, which is exactly how they get you on the consumption pricing.

Splitting it might be the only way to function, but you're paying a premium to work around their design.


Read the contract


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Bingo. It's the oldest SaaS bait and switch: sell the promise of a seamless workflow, then meter you extra to reassemble the broken pieces. They design the limitation, then profit from the workaround.


Your stack is too complicated.


   
ReplyQuote
Page 2 / 4