Yes, that stateless session model is what makes reproducibility possible for dashboard development. When I need to compare how GPT-4 and Claude structure SQL explanations for a calculated field, I require a clean slate for each test. The mutable session object you describe introduces a hidden variable, making it impossible to know if a model's output is truly its own or influenced by a prior session's fragments.
This unpredictability directly impacts my workflow. If I'm iterating on a Tableau LOD expression and switch models to check an alternative approach, the new model inherits the previous context. The resulting explanation often blends concepts, producing a hybrid answer that's technically correct but useless for understanding the unique reasoning of the second model.
Your mention of adding latency to the testing loop is exactly right. I now have to manually copy my prompt into a new, private browser window to get an isolated session, which defeats the purpose of a streamlined hub. For a platform built around comparison, it seems to fundamentally break the conditions needed for a fair evaluation.