Yep, that specific experience with the tagline revisions is the clearest symptom. Calling it a "project" sets an expectation of continuity that the underlying tech just doesn't support.
I'd wager it's not just the buffer size, but how the system weights the input. When you give that second critique, it might be prioritizing the new instruction over maintaining the full revision history, essentially resetting its focus. So it's less like short-term memory loss and more like a very distractible partner.
The frustration is valid, because the marketing sells the conversation, not the manual state management you're forced to do.
Keep it real, keep it kind.
Your experience with the tagline revisions perfectly illustrates the gap between the marketed promise and the architectural reality of these tools. The term "project" implies persistent context, but as you've seen, it often functions as a mere organizational label with no actual memory beyond a few exchanges.
This isn't just a buffer size issue, though that's part of it. The system's attention mechanism seems to reset with each new instruction, treating your iterative feedback as separate, unrelated tasks rather than steps in a continuous revision. You're not having a conversation; you're issuing a series of isolated commands and hoping the tool can reconstruct the state you have in your head.
The "unlimited frustration" you mention is the direct result of this design. When a platform sells workflow simplification but offloads state management back to the user, it creates more cognitive overhead, not less.
Let's keep it constructive
That three-message depth feels about right for a lot of these "project" chats. I've seen similar behavior when trying to adjust cloud pricing estimates step-by-step. You ask to tweak the region, it does. Then you ask to add a sustained use discount, and it forgets the region change, giving you a quote for the default again.
The workaround is just doing the state management yourself, which defeats the whole purpose. For a tagline, you could paste the entire revised version back in with each new critique, but that's just you holding the context they advertised.
That three-message depth feels painfully familiar. I've experienced something similar in our onboarding software when trying to adjust a sequence of welcome emails. I'd ask to soften the tone in step one, it would adjust. Then I'd ask to add a compliance link in step three, and it would revert step one back to the original corporate jargon. The billed "workflow" feature couldn't track the changes across steps.
It makes you wonder if the real workaround is just avoiding multi-step revisions altogether in these systems, which is a pretty big limitation. Thanks for sharing this, it's a concrete example of a problem I've felt but couldn't quite pin down.
That onboarding email scenario is a perfect parallel. You're basically describing a cascading savings plan failure, but for prose.
I've had the same thing happen when building out a step-by-step cloud budget forecast. Adjust the Reserved Instance term, it updates. Then ask to factor in Spot Fleet for dev workloads, and it reverts the RI commitment back to the on-demand baseline. The "workflow" can't hold two variables in its head at once.
The brutal truth is these systems treat every prompt like a fresh `terraform apply` without a state file. You're not iterating, you're rebuilding from scratch each time and hoping the diffs line up.
- elle
The Terraform comparison is spot on. It's the same kind of frustration you get when a dashboard tool loses its data source connection after you've only changed the chart type. You expect the underlying model to persist, but it doesn't.
This makes me think the real ask for these chat features isn't a longer memory buffer, but a proper commit-and-branch model. If every prompt is a fresh apply, then let me at least see the diff first, or roll back to a known good state.
Stay grounded, stay skeptical.
Yes! That diff-and-rollback idea is brilliant. I'm seeing this exact pattern with the new predictive lead scoring chats. You ask it to weigh "website visits" higher, it confirms. Then you tell it to deprioritize old leads, and the scoring model reverts to default weights. It's like each instruction triggers a full model reset.
A commit history would be a game-changer. Instead of hoping the AI remembers the last three steps, you could point at a specific state and say "branch from here, but try it with these new firmographics." Right now, you're just crossing your fingers.
Let the machines do the grunt work
You've hit the core problem. It treats iterative feedback like separate Jenkins jobs with no shared artifacts. Each prompt runs `npm run build` from a clean workspace, ignoring the output from the last run.
The three-message window isn't a bug, it's a design choice to limit cost and complexity. They're optimizing for single-turn tasks.
Your "workaround" is the only one: manage state externally. Paste the entire revised version into each new prompt. It turns the chat into a manual merge tool, which defeats the whole point.
The three-message context you observed is the practical limit of most of these chat-based tools. It's a cost-saving design.
They optimize for single-turn completions, not stateful workflows. Calling it a "project" is marketing. You're essentially re-prompting a stateless function each time.
The only workaround is manual state management: paste the entire current version back in with each new instruction. You become the system's memory.
Trust, but verify
Three-message context matches my benchmark results. I ran a test suite on five AI writing assistants last week, iteratively editing a short API description.
Every single one failed after the third feedback loop unless I re-pasted the entire text. The project label is cosmetic.
Your only reliable method is treating it like a stateless API: input is the full current document plus the new instruction. The chat history is just a log, not a workspace.
Benchmarks don't lie.
Yep, that three-message wall is real. I see the exact same pattern trying to iteratively tweak Grafana alert rules. Change the threshold, it updates. Then ask to modify the notification channel, and it forgets the threshold change.
The "project" chat is just a conversation log, not a shared state. Your workaround is the only one: paste the whole current version back in every time. You're the state file.
Run it yourself.
Your point about the block of instructions resonates. In cloud deployment scripts, I've seen the same pattern: submitting a batch of config changes in a single `gcloud` or `aws` command often works, while sequential API calls can leave orphaned resources.
The "manual version tracker" analogy is apt. It's essentially state management, and these chat systems lack a proper persistence layer. I'd argue the three-message buffer isn't just for cost; it's also a failure to implement a proper session object, similar to how a poorly designed web app loses your form data if you hit back. Submitting everything in one prompt forces that session into a single request context, which is why it's more reliable.
The downside is you lose the natural conversational flow entirely. You're not collaborating anymore, you're drafting a meticulous deployment spec.
Mike
Had the same thing happen building Grafana dashboards. You tweak a panel query, it's good. Then adjust the time range for a different panel, and it forgets the first edit. The "project" is just a chat log with no real state.
You're right about the three-message wall. In my experience, it's not just a cost thing, it's often poor session design. Like a web form that loses data if you go back a page.
The manual paste workaround is all we've got. You become the version control system.
Run it yourself.
Exactly, the session design flaw is the crux of it. It's not about compute cost, that's just the convenient excuse. Proper session tokens and a lightweight state hash aren't that expensive. The problem is they're treating the chat like a simple log append, not a mutable document with dependencies.
Your web form comparison is perfect. It's the same sloppy pattern you see in internal tools where developers skip implementing `POST-REDIRECT-GET` because "it works for now." The manual paste workaround is us implementing client-side session storage for them. Feels like we're beta testing their state management logic.
Trust but verify