Okay, so I was deep in a proof-of-concept for an internal support bot, testing two different AI models for generating SQL from natural language questions. The usual drill: build two separate endpoints, orchestrate the calls, log everything, compare outputs... a real pain.
Then a colleague pointed me to a feature in **Claw** (the LLM gateway we're evaluating) called **dual-run mode**. It's shockingly simple. You define your two agents (different models, prompts, parametersβwhatever) and call them in a single request. It returns both completions side-by-side with full metadata.
Hereβs a bare-bones config example from my test:
```yaml
agents:
agent_slow_but_smart:
model: "gpt-4"
system_prompt: "You are a precise data analyst. Generate ANSI SQL only."
temperature: 0.1
agent_fast_and_good:
model: "claude-3-haiku"
system_prompt: "You are a helpful data assistant. Generate ANSI SQL only."
temperature: 0.2
dual_run:
agents: ["agent_slow_but_smart", "agent_fast_and_good"]
request: "Get total sales last quarter by region"
```
The response bundles both outputs, with latencies, token counts, and the full messages. I dumped the results into a small Looker dashboard to compare accuracy and speed over hundreds of test queries. Made the decision data-driven instead of gut-feel.
Has anyone else used a similar side-by-side capability in other LLM orchestration tools (like LangChain, etc.)? I'm curious about:
* How you handled the evaluation metrics (beyond just eyeballing the SQL).
* If you ran this in production during a cutover, how you routed traffic (e.g., % split) based on the dual-run results.
* Whether you kept the dual-run pattern live for ongoing model monitoring.
The real win for our migration was killing the "my bot vs. your bot" debate with hard numbers. 😄
--diver
Data is the new oil - but it's usually crude.
Wait, so does Claw handle the cost tracking for each agent separately in this mode? That's a huge pain to do manually when you're trying to compare not just performance but also price per call.
Still learning
Yes, it does track cost separately. The response object includes a breakdown of usage and estimated cost per agent, which you can pipe directly into a spreadsheet. The real time-saver is that you don't have to reconcile separate billing logs from two API providers. Makes it a lot easier to decide if a 2% accuracy boost is worth a 300% cost increase.
Connecting the dots.
Wow, that sounds like a huge time saver. I've only ever used Asana and ClickUp, so this whole LLM gateway thing is new to me. When you say it bundles the outputs, does it also flag if the generated SQL is syntactically different? Or is it just a raw text comparison you have to do yourself?