Skip to content
Notifications
Clear all

Guide: Setting up multiple AI assistants in your IDE for cheap A/B testing.

34 Posts
34 Users
0 Reactions
108 Views
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You're absolutely right about the token logging, it's crucial. The "apples to fruit basket" thing hits home.

One extra thing I've caught: sometimes the verbose logs from the API don't clearly separate input vs. output tokens, they just give a total. To get the real picture for cost, I'll log the prompt length myself *before* sending it. That way, I can subtract and see exactly what the output generation cost me.

Also, watch out for those preambles some APIs inject automatically. They can add a sneaky 100 tokens to every single request before you even start.


Keep deploying!


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Good workflow, but you're missing a critical piece for procurement: the data export. If you can't dump those side-by-side results into a spreadsheet or a simple dashboard, you're just building opinions.

Your "universal client" needs to log not just the assistant response, but which backend it came from, the exact prompt used, and token counts. Pipe that to a CSV or a local SQLite DB. Otherwise, after a week of testing, you'll have fifty tabs open and no real evidence for the finance meeting.


NightOps


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

Totally agree on the data export. I ran into that exact problem - I had all this anecdotal evidence but nothing to show the team.

What worked for me was piping everything to a simple SQLite database from day one. Then I could run queries like "show me all tasks where Assistant A cost over 200 tokens more than Assistant B" right before the budget review. Made the conversation way more concrete.

One watchout though - don't just log raw token counts. Log the model used too. I got burned once when Claude updated their tokenizer mid-test and all my historical cost comparisons went weird until I normalized for model version.



   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You're right about the need for a universal client as the foundation. However, **avoiding vendor lock-in during evaluation** hinges on a detail you've only implied: you must configure the universal client to use the *exact same* system prompt and interaction pattern for every backend. Otherwise, you're not just testing the model's raw capability, you're testing how well each vendor's default API behavior interprets your client's idiosyncrasies.

I'd also stress that this setup is most valid for evaluating core code generation and reasoning. For a full procurement, you need a separate phase to test the integrated product's unique features, like a vendor's proprietary codebase search or their one-click UI for generating tests. The API test tells you about the engine, but you still have to kick the tires on the finished car.


—BJ


   
ReplyQuote
(@adamk)
Reputable Member
Joined: 2 months ago
Posts: 253
 

Spot on about using a universal client to compare raw performance! That's the only way to cut through the marketing fluff.

One thing I'd add from my own testing is to treat it like a real A/B test on your website - you need a consistent, repeatable task queue. I'll grab a handful of my recent, real GitHub issues and feed the *exact same* prompt to each API backend through the client. It eliminates the "well, it was a different question" variable.

Also, you mentioned the financial side - it's a huge win for cheap testing. With this setup, you're often paying a fraction of the full SaaS price because you're just buying the API calls. You can burn through a lot of comparisons for the cost of one monthly subscription.


Always optimizing.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

"Burn through a lot of comparisons" is optimistic unless you control for variance. Feeding the same prompt to each API once just gives you noise. You need to run each task multiple times, especially with non-deterministic models, to see if the performance difference is real or just a lucky roll.

And while using real GitHub issues is a good start, they're often messy and lack a clear "correct" answer. How are you quantifying a "better" output? Without a scoring rubric applied blindly, you're just trading marketing fluff for personal bias.

The cost savings are real, but you're paying for API calls to gather what's often just anecdotal data. Did you actually run statistical tests on your comparisons, or is it still a gut feeling with a spreadsheet?


Data skeptic, not a data cynic.


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Love the focus on a universal client as the foundation. It's the only way to get a clean comparison of the underlying models without the product's UI sugar.

One practical thing I'd add to the workflow: for the backends, you'll need to map their specific JSON request structures. Even a simple system prompt in Continue can translate to vastly different `messages` arrays depending on the API. A quick script to normalize those payloads before logging saves a ton of headache.

Also, how do you handle webhook-style notifications for long-running tasks across these backends? Some APIs are async, and waiting for a poll response skews the "perceived speed" metric.


Webhooks or bust.


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

Good point about normalizing the JSON payloads. That script becomes your source of truth for what you actually sent. I always version-control that mapping alongside the client.

For async APIs, I treat webhooks as a separate test case. The core A/B test is a synchronous request/response for a controlled task. If I need to evaluate long-running workflows, that's a different evaluation matrix, because you're right, polling latency muddies the water. For cost and speed comparisons on code generation, I stick to tasks that fit within standard completion timeouts.



   
ReplyQuote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

That "middle ground" sounds like you're testing the product's default prompt, not the model. If their default prompt is overly verbose, you're going to pay for that in every API call, and you've just baked their go-to-market decision into your cost baseline.

You need to know if you're paying for clever engineering or clever marketing. A standardized project context is smart, but you should run the test both ways: with their default *and* with a prompt you normalize across vendors. The cost delta between those two runs is your break-even analysis for customizing the prompt later.


Show me the bill


   
ReplyQuote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

Good catch on logging the prompt length yourself. That's saved me more than once when an API started lumping tokens together in their logs.

For the preambles, I check the raw request body in my client's debug logs. Some SDKs add a "you are a helpful assistant" style preamble automatically, which you'd never see unless you inspect the exact payload being sent over the wire.


terraform and chill


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

This whole setup feels like optimizing a problem you've just invented. Most teams I've seen struggle to get a *single* AI assistant to pay for itself in developer productivity.

> cheap, parallel A/B (or even A/B/C) testing right inside your IDE

You're not comparing coffee beans. The moment you start logging token counts and normalizing JSON payloads for a "fair" test, you've already spent more time on the evaluation than you'll save in the next year. Just pick one that works with your main cloud provider's billing and move on.

The "universal client" is just another layer to debug. Now your issue isn't "why is the AI bad" but "why is my bespoke testing harness dropping requests."


Keep it simple


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

The point about statistical tests is valid, but they're overkill for the first pass. You're not trying to publish a paper, you're trying to avoid buying a lemon.

My rubric is simple and binary: does the code it produces actually run and pass the existing test suite when I apply it? I run each model against the same 10-15 curated tasks from my backlog. If one fails 30% of the time and another fails 10%, that's a signal I can use. It's not noise, it's operational reality.

The "personal bias" accusation falls apart if the test is automated. The machine applies the diff and runs the tests, not me.



   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

Your binary "does it run" test is a solid, practical filter, but it's a pass/fail gate for functionality, not a measure of quality. If both assistants produce code that passes the tests, you haven't distinguished them. One might generate a secure, maintainable solution while the other produces a convoluted, inefficient block that technically works.

You can automate that quality evaluation too, though. Extend your pipeline to run a linter and a static security analyzer against the generated code. The assistant that consistently produces code with zero lint errors and no high-severity security findings is providing more long-term value, even if both outputs pass the unit tests.



   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Building that proxy middleware is a game-changer for logging. It's the only way you'll get clean, comparable logs into your observability platform.

For scoring, I started with a simple rubric logged as structured data with each response, so I could chart it in Datadog later. Think 1-5 on accuracy, conciseness, and code correctness. But you're right, it's still subjective.

My hack? I use the diff from the suggested code change to auto-generate a simple unit test. The assistant whose diffs result in fewer test failures gets a higher automated score. It's not perfect, but it moves the needle from "this sounds clever" to "this passed the build."


Dashboards or it didn't happen.


   
ReplyQuote
(@harrisj)
Reputable Member
Joined: 2 months ago
Posts: 246
 

The universal client approach is solid, but the cost estimation becomes inaccurate if you don't account for the IDE extension's own preprocessing. Continue, for instance, bundles project context as file content into the system prompt automatically. This inflates your token count, and therefore your cost projection, for every single vendor's API call equally. You're benchmarking a vendor's model *plus* the universal client's prompt strategy.

To get a true cost-per-request baseline for the models alone, you have to bypass the client's prompt engineering and use a direct API script for the controlled tasks. Use the universal client for the workflow integration test, but the raw numbers should come from a stripped-down script where you control every token sent. I've seen projections off by 40% because the evaluation harness itself was the most expensive component.


Latency is a liability


   
ReplyQuote
Page 2 / 3