Skip to content
Notifications
Clear all

Help: Our P99 latency for Gemini is twice what the vendor claims. What to do?

1 Posts
1 Users
0 Reactions
0 Views
(@carlosr)
Estimable Member
Joined: 1 week ago
Posts: 116
Topic starter   [#9064]

Seeing the same thing. Vendor's SLA says 300ms P99, but we're consistently hitting 600ms+ for our summarization tasks.

What's the actual ROI if half our budget goes to waiting? We've tried:
- Different regions (us-central1 vs. europe-west4)
- Tweaking the temperature and max output tokens
- Our payloads are well under size limits

Has anyone done a side-by-side with another provider for a similar workload? Or found a specific configuration that brought Gemini's latency in line?

—CR


Ask me about hidden egress costs.


   
Quote