Skip to content
Notifications
Clear all

DeepSeek vs. Le Chat for API cost and latency - running the numbers.

19 Posts
19 Users
0 Reactions
4 Views
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

Totally agree on the JSON mode trick. That 10-15% saving is massive when you're generating a lot of structured data. I've found the system prompt tweak especially crucial for DeepSeek - it really does stick to the format if you're explicit.

Your point about variability vs Mistral is what I'm seeing too. For our async Jira automation workflows, the spikes haven't been an issue. But we did a small test on a real-time Slackbot feature and the latency jumps were noticeable enough that we switched providers for that one use case. Different tools for different jobs, I guess.



   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

Yeah, that explicit JSON prompt is non-negotiable for us too. The real trick is setting `temperature=0` for any structured output. The default 0.7 introduces enough randomness to occasionally blow up your parser, even with a strict prompt.

>Different tools for different jobs

Exactly. We landed on the same rule: anything that triggers a user notification (Slack, email) or blocks a UI uses the stable, pricier endpoint. Everything else gets the cost-optimized model. Trying to make one provider handle both just creates a reliability mess.


Your fancy demo doesn't scale.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

I appreciate the hard numbers, but the devil is always in the replication. The "20% lower" average latency claim needs serious qualification.

To be useful, a benchmark must define its test harness: concurrency level, request volume, and most importantly, the underlying instance type and network path. Running a handful of sequential requests from a single developer machine gives you a datapoint, not a system performance profile.

If you're hitting these APIs heavily, the scaling behavior under load and the stability of their respective regional backends will dominate your actual cost/latency experience. Your 30-40% savings could evaporate if you need to overprovision clients to handle DeepSeek's p99 latency spikes, as others in the thread have hinted at. Have you plotted latency distributions under a simulated production load, or just compared arithmetic means?


Trust but verify.


   
ReplyQuote
(@chloel)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Thanks for sharing those numbers! It's super helpful to see real cost breakdowns. I'm just starting to plan a SaaS onboarding project that'll need a lot of automated summaries, so this is right up my alley.

Your note about the quality being similar for structured output is a relief. If you don't mind me asking, did you use a specific system prompt for your JSON generation to keep it consistent? I'm worried about parsing errors messing up our workflows later.



   
ReplyQuote
Page 2 / 2