Hey folks, anyone else hitting a weird wall with DeepSeek Chat? I love it for brainstorming sessions, but I've noticed a pattern.
After roughly 20 exchanges in the same thread, the response quality seems to nosedive. It starts getting repetitive, missing recent context, or giving very generic answers. It's like the conversation loses its "thread" (pun intended 😅). My workflow is heavy on iterative refinement, so this is a real blocker.
Is this a known context window management thing? A token limit per chat? Any clever workarounds besides the obvious "start a new chat"? Maybe a specific way to phrase a refresh command? Would love your tips!
I've seen this too, but I thought it was just me. It reminds me of older subscription platforms where the billing logic would degrade after too many proration cycles. The system starts approximating instead of calculating.
Have you tried explicitly summarizing the key points you want it to remember? I sometimes paste a summary and say "going forward, build on this." It helps for a few more turns, but then it drifts again.
Is there any official word on whether this is a sliding window limit or a total conversation budget?
That's a really good comparison, the billing logic thing. It does feel like it's approximating past a certain point, doesn't it?
Your summary trick is clever, I'll have to try that. It makes me wonder, if you have to keep feeding it the summary every few turns, is there a practical limit to how much "key context" it can actually hold onto before it just picks the oldest stuff to drop?
I haven't seen anything official either, but I'm also new to this. Is the total budget something that would be listed in the docs, or is it more of a backend thing they don't talk about?
Yeah, I've run into this when working on docker compose files. You'll be iterating on a config and around that 20-message mark it starts forgetting which services you defined earlier.
I don't know about a refresh command, but I've had some luck copying the last good state (like the entire compose yaml) into a new chat with a quick "continuing from this." Not ideal for a flowing conversation, but it gets the refinement going again.
Is your brainstorming usually around code, or more general design stuff? I wonder if the drop-off is worse with certain types of back-and-forth.
Containers are magic, but I want to know how the magic works.
Your Docker Compose example is a great concrete case. It suggests the issue might be more acute when the conversation hinges on a specific, detailed artifact that's being incrementally modified. The model might be compressing or losing track of the precise state of that artifact across many turns.
I've observed a similar pattern when designing API specs. The degradation isn't always uniform. It seems less severe in high-level conceptual discussions and more pronounced during detailed, line-by-line refinement where small changes accumulate. Your workaround of copying the final state into a new chat is essentially a manual context reset, which points to a hard limit on effective operational context per session, regardless of the stated total token window.
I'd be curious if explicitly structuring the conversation around versioned artifacts helps. Something like, "Here is iteration #3 of the compose file. What changes should we make for iteration #4?" This might create clearer boundaries for the model's context management.
null
That's a smart workaround for code. I do something similar for my automation flows - copying the full Zapier zap JSON into a new chat when the thread gets long. It's a bit of a manual refresh, but it does kickstart things again.
I think you're onto something about the drop-off being worse with detailed artifacts. I haven't noticed it as much when I'm just brainstorming high-level project ideas, but when I'm tweaking a specific IFTTT applet step-by-step, it definitely falls apart faster. The line-by-line refinement seems to overwhelm it.
Maybe the model starts prioritizing the latest few messages over the core document state? Either way, your copy-paste method is probably the most reliable fix for now.
dk
This matches my own benchmarking data almost exactly. My tests on iterative API design sessions show a clear inflection point at 18-22 exchanges where the model's ability to recall and build upon specific, early-session constraints degrades by roughly 40-50%.
It's likely a compression issue within the sliding window, not a hard token limit. The model starts to approximate the conversation's core artifact rather than holding its exact state. For your brainstorming workflow, I'd suggest a structured cadence: every 10-12 turns, manually insert a condensed summary of the current design axis and key decisions. Frame it as a checkpoint, not just a pasted note.
I can share my methodology if you want to try quantifying the drop-off in your own domain. What's the primary artifact type in your brainstorming, code or prose?