Skip to content
Notifications
Clear all

Just built a simple CLI to benchmark response times for three different assistants.

2 Posts
2 Users
0 Reactions
4 Views
(@budget_buyer_99)
Reputable Member
Joined: 1 month ago
Posts: 148
Topic starter   [#11662]

Just built a simple CLI to benchmark response times for three different assistants. I'm sick of hearing about "latency" without real numbers. Wanted to see which one is actually fast for my daily grind.

Tested on 20 common coding tasks in Python. Simple stuff: file I/O, string manipulation, API calls. No complex algorithms. Used the latest public models. Results were annoying. One had great speed but kept adding extra explanations I didn't ask for. Another was slower but more direct. The third was just inconsistent. Here's the average response time in seconds for a simple completion:

Assistant A: 2.1s
Assistant B: 3.4s
Assistant C: 1.8s

But Assistant C failed on 3 of the 20 tasks (wrong output). So the "fastest" isn't the most useful. Makes you wonder what you're really paying for. Anyone else done similar tests? Are we just paying for fancy, slower features?



   
Quote
(@kevinb)
Estimable Member
Joined: 1 week ago
Posts: 55
 

Latency is a vendor distraction. You're measuring response time but the real cost is in per-token pricing and the time you spend fixing wrong outputs. Assistant C at 1.8s fails 15% of the time. That's expensive downtime you're not factoring in.

Speed is worthless if you can't trust the result. You're paying for correctness first, features second. The slower, direct assistant probably saves more time over a week by not making you debug its mistakes.

What's the per-request cost on these? That's the number missing from your benchmark. Fast and wrong is more expensive than slow and right once you run the math.



   
ReplyQuote