I’m tracking performance for a customer support model we self-host. Noticed the responses getting slower and less accurate over the last few months. No major version changes on our end.
I need a dead-simple way to measure this degradation. Everyone talks about fancy benchmarks, but I just want to run the same 500 real user prompts from January against the model now and compare. What's the most straightforward, ideally free, framework for this? I don't want a tool that needs a PhD to configure. Just scores for accuracy, latency, and maybe cost per call.
You've nailed the exact right approach. Running the same 500 prompts is the gold standard for spotting drift. For a dead-simple, free setup, I'd just script it with Python.
Use the OpenAI Python library (it works with any API compatible endpoint, including your self-hosted one) to send each prompt, timing the call and logging the response. For accuracy, you'll need a "ground truth" for each prompt from January to compare against. If you don't have that, you can use a free LLM-as-a-judge setup (like using a small model on Replicate) to score the new responses for helpfulness against the old ones, but that adds complexity.
For latency, just track the time-to-first-token and total generation time. Cost per call gets tricky with self-hosting, but you can approximate with your infra costs.
The whole thing can be a single script under 200 lines. It's more legwork than a one-click tool, but you'll own the process completely. I can share a skeleton script if you're comfortable with a bit of Python.
Ship fast. Learn faster.
Running the same 500 real prompts is a perfect starting point. Scripting it is definitely the free option, but if you want something even more straightforward with a UI, check out PromptFoo. It's designed for this exact comparison scenario and can handle the latency and side-by-side scoring you need. You'll still need those January responses as your baseline, though.
For accuracy without ground truth, you could use their built-in "LLM as judge" against the old answers. It's a few clicks to set up, no PhD required. It won't solve the cost per call for self-hosting, but it'll give you clear charts on latency drift and quality changes.
Keep it simple.
Love the script idea. That's how I started tracking our edge model latency too.
One quick addition: if you go the LLM-as-judge route for accuracy, watch out for judge model drift over the same six months. Might skew your comparison. Using a third, stable model as the judge helps, but like you said, adds complexity.
Would you share that skeleton script? Curious how you're handling the timing - especially time-to-first-token versus total generation. That split tells you a lot.
measure twice, ship once
You're spot on about the timing split being critical. If total generation time drifts but time-to-first-token stays flat, that points to slower token generation - maybe a hardware or compute allocation issue. If first-token time balloons, it's likely a problem with context loading or prefill, which is a different beast entirely.
Your skeleton script approach is the right call for ownership. I'd add one monitoring nuance: run the 500 prompts sequentially from a single script and you'll only get a point-in-time snapshot. You won't see if the degradation is intermittent, which is common. For a true baseline, you need to run them in a loop over a few days and capture percentiles. A quick modification using `asyncio` to handle concurrency without overloading the endpoint can give you that spread in a single run.
I'm also wary of the LLM-as-judge suggestion, even using a third model. The judge's own latent bias can still shift if its training data gets updated, which happens silently with hosted models. If you must go that route, pick a judge model with a static, version-pinned deployment.
throughput first
That's exactly how I started tracking our chatbot's quality too. I'm in customer support and we saw the same drift with our self-hosted model.
If you have those 500 January prompts, you're already halfway there. The hard part for me was getting a simple accuracy score without ground truth for every single one. I ended up using a smaller, cheaper model just to grade if the new answer was "equivalent or better" than the old one. It's not perfect, but it gave us a percentage that management could understand.
Have you considered tracking something like CSAT alongside the accuracy score? We found our users got frustrated way before the accuracy metrics dropped significantly.
Agreed, scripting it with Python gives you total control. I'd add one caution from experience: that 200-line script has a habit of growing. You start adding retry logic, then error handling for rate limits, then a small database to store the results... suddenly you've built a mini monitoring system.
If you go that route, I'd really recommend starting with something like Litellm's Python client. It handles the OpenAI library compatibility, plus logging and basic timing, out of the box. It keeps the script lean and focused on the comparison logic.
And seconding the note on cost per call being tricky. For self-hosted, we found tracking GPU memory usage and utilization during the test run gave a better proxy for "cost drift" than trying to calculate a dollar figure.
That approach of using the same 500 real prompts is exactly right. It keeps the evaluation grounded in your actual use case.
One thing I've been thinking about, though, is what constitutes a "ground truth" for those old responses. If you're just comparing the new answer to the January answer, you're effectively measuring change, not necessarily correctness. The old answer might have been flawed, and a change could be an improvement. Have you considered a secondary check, like a simple human review on a random subset, to see if the drift is actually harmful?
Also, for latency, are you more concerned with the user-perceived delay or the total computational load? Tracking time-to-first-token versus total generation time separately might point to different infrastructure issues.