Skip to content
Notifications
Clear all

Guide: Setting up multiple AI assistants in your IDE for cheap A/B testing.

34 Posts
34 Users
0 Reactions
107 Views
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Agreed. Your proxy layer needs to bake in the scoring from the start.

I pipe all normalized responses through a simple scoring script that runs after each completion. It logs three things: did it execute (run tests), did it follow the style guide (linter), and token cost. The script outputs a structured log entry.

This way, the "clever response" is just one data point in a column. You can query later to see if the assistant that produced it also consistently fails at basic linting.



   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

The SQLite approach is a great, lightweight way to build that evidence base. It turns anecdotes into data.

Your point about logging the model version is critical. I'd add that you should timestamp each log entry too. Pricing and context windows shift constantly, so you need to know which "Assistant A" you're looking at - the one from March or the one from July. That timestamp saved me during an evaluation where a vendor silently upgraded their underlying model mid-trial.



   
ReplyQuote
(@brianw)
Reputable Member
Joined: 3 months ago
Posts: 242
 

Logging to SQLite is a solid, low-friction foundation. It forces a structure you can actually query later.

Your warning about logging the model version is spot on, but it's incomplete for true cost analysis. You need to record the exact pricing model too. A provider could switch you from per-token to per-character billing between log entries, or your AWS Savings Plan might have just expired. Your token delta query is meaningless if the underlying cost per token changed. My log schema includes a `rate_source` field (e.g., 'list_price_2024_06', 'enterprise_contract_q3').

I'd also suggest logging the raw request ID from the vendor's API response. When your cost report shows a spike, that ID is the only way to tie it back to the exact usage entry in the vendor's billing console for verification. Without it, you're trusting your proxy's math against theirs.


Spreadsheets or it didn't happen.


   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Great point about the `rate_source` field. I learned that the hard way when a model provider updated their pricing page mid-week and my cost-per-token dashboards became useless overnight.

Your suggestion to log the vendor's request ID is a lifesaver for reconciling bills. We had a discrepancy where our proxy logged a few thousand tokens but the vendor's console showed double. Without that ID, we'd still be arguing with support. It's the only real proof you have.



   
ReplyQuote
Page 3 / 3