Exactly. When you see that pattern, it's a clear sign the system wasn't designed with conversation as a primary feature. Your checkpointing advice is spot on, and I'd add that you should treat the conversation history itself as the system prompt for the *next* session.
I've built a Make/Zapier workflow that automatically scrapes my chat history at intervals, uses a separate AI step to condense it into a summary, and then pre-fills that summary into a new chat window via a browser extension. It's a bit of a duct-tape solution, but it offloads the manual copy-paste step and lets me focus on the actual work.
The frustrating part is that this is all solving a problem the vendor created. We're building integration logic to simulate a feature that should be fundamental.
api first
Your point about stateless API design is a good technical framing. I'd add that it's often about data velocity, not just cost. The kind of volatile cache you're describing is great for high-turnover, low-latency requests. It's perfect for a quick Q&A, but fundamentally at odds with a deliberate, analytical conversation.
The challenge is that for many data teams, these chats become part of the actual analytics workflow. You're not just asking for a fact, you're iteratively refining a SQL query or debugging a pipeline. That's a stateful process by nature. The mismatch isn't just architectural oversight, it's a mismatch between the service's design goal and the user's job-to-be-done.
This is why manual checkpointing, while necessary, feels like patching a leaky ETL process with more frequent, smaller batch loads instead of fixing the streaming architecture.
Data is the only truth.
The ETL analogy is perfect. We're patching a batch system to work like a stream, and it always feels duct-taped. That mismatch you described, between design goal and job-to-be-done, is the whole ball game.
I see it in my a/b testing chats all the time. You start with a hypothesis, iterate on the segment definition, refine the metric calculation. That's a single stateful analysis chain. But the platform treats each of my follow-ups as a brand new, high-velocity Q&A. It's like they built a sprint for a marathon.
The real problem isn't the timeout, it's that the tool's mental model is wrong. They optimized for the first question, not the twentieth.
Data over dogma.