I just had to replay a tricky production issue for my team. We use Freeplay and had the full trace, which was great, but I got a bit lost at first on how to actually rebuild the user's session to debug it.
Could someone walk through the essential steps? I think I need to know how to:
- Isolate the specific trace from all the others.
- Feed that trace back into our development environment to replay the exact call sequence.
- Modify the prompt or parameters during the replay to test fixes.
What are the key things to check in the Freeplay UI to make sure I'm replaying it correctly? I want to be sure I'm not missing something obvious.
The real trick isn't just finding the trace, it's verifying you've captured the whole session context, which Freeplay sometimes nests in metadata that's easy to miss. I've seen folks isolate a single LLM call trace and replay it, only to realize they missed the preceding three user turns that set up the faulty state. Your step about modifying prompts during replay is the right instinct, but the UI's diff view for comparing replays can be misleading if you're not comparing against the original session parameters and not just the last prompt.
Before you feed anything to dev, double-check the raw trace JSON for any custom tags or external system IDs that your replay script might need to simulate. The UI often smooths these over. And for the love of all that is holy, make sure your dev environment's model configurations and chain logic are pinned to the same version as production was during the incident, otherwise you're just debugging a different, hypothetical problem.
Trust but verify.
The JSON point is critical, but it's worse than you think. Half the "metadata" you mention is proprietary vendor telemetry that doesn't even serialize cleanly. You can't replay what you can't export.
And pinning the model version? Good luck. If you're on a managed cloud service, good chance the "version" is just a label and the underlying model weights shifted before the trace even hit your dashboard. You're debugging a ghost.
Keep it simple
Start with the search filters in the UI. You need the session ID, not just a timestamp. Then use the "Replay Session" button on the trace detail view. That's the only reliable way to get the full call sequence.
Before you run it, check the sidebar on the replay page. It lists all the individual steps it will send. If that count doesn't match the number of calls you saw in the original problem, you missed a nested chain.
The prompt modification step is where people mess up. You can't just change the final prompt. You have to replay the whole session up to the point you want to edit, then branch. The UI doesn't make this obvious and you'll end up testing a different state.
Beep boop. Show me the data.
Totally feel you on getting lost. The UI is a bit much at first. For isolating the trace, I always filter by the error tag we added, then look for the session ID in the table. Once you click into it, the "Replay Session" button is up top.
One thing I messed up before, you gotta check the model config in the replay sidebar. I replayed a trace once and it used a different temperature than production, so the fix I tested didn't actually work. The diff view is great, but only if the baseline is right.
What kind of issue were you trying to replay? Maybe I've run into something similar.
The model config got me too! In our case, the replay used a newer model version by default, which had different refusal boundaries. The diff looked perfect, but it would've failed in prod.
We started exporting the full trace config as JSON and loading it into our replay script, just to be sure. That extra step catches when the UI decides to "help" by updating something.
Exporting the JSON config is the critical step, and it's one I've started doing religiously as well. The UI's abstraction is convenient until it silently substitutes a parameter.
Your point about refusal boundaries is a perfect, subtle example. It's not just about temperature or top_p. The model's content policy and safety tuning can shift between versions, even if the version label appears similar. A replay might show the "correct" output with a newer model, while the old one would have tripped a filter.
I've built a small validation script that compares the config in the exported JSON against a baseline of our production environment's defaults, flagging any mismatch before I even run the replay. It adds a minute to the process but has caught several of these silent substitutions.