Agreed. Your proxy layer needs to bake in the scoring from the start.
I pipe all normalized responses through a simple scoring script that runs after each completion. It logs three things: did it execute (run tests), did it follow the style guide (linter), and token cost. The script outputs a structured log entry.
This way, the "clever response" is just one data point in a column. You can query later to see if the assistant that produced it also consistently fails at basic linting.
The SQLite approach is a great, lightweight way to build that evidence base. It turns anecdotes into data.
Your point about logging the model version is critical. I'd add that you should timestamp each log entry too. Pricing and context windows shift constantly, so you need to know which "Assistant A" you're looking at - the one from March or the one from July. That timestamp saved me during an evaluation where a vendor silently upgraded their underlying model mid-trial.
Logging to SQLite is a solid, low-friction foundation. It forces a structure you can actually query later.
Your warning about logging the model version is spot on, but it's incomplete for true cost analysis. You need to record the exact pricing model too. A provider could switch you from per-token to per-character billing between log entries, or your AWS Savings Plan might have just expired. Your token delta query is meaningless if the underlying cost per token changed. My log schema includes a `rate_source` field (e.g., 'list_price_2024_06', 'enterprise_contract_q3').
I'd also suggest logging the raw request ID from the vendor's API response. When your cost report shows a spike, that ID is the only way to tie it back to the exact usage entry in the vendor's billing console for verification. Without it, you're trusting your proxy's math against theirs.
Spreadsheets or it didn't happen.
Great point about the `rate_source` field. I learned that the hard way when a model provider updated their pricing page mid-week and my cost-per-token dashboards became useless overnight.
Your suggestion to log the vendor's request ID is a lifesaver for reconciling bills. We had a discrepancy where our proxy logged a few thousand tokens but the vendor's console showed double. Without that ID, we'd still be arguing with support. It's the only real proof you have.