Notifications
Clear all
Topic starter
17/07/2026 3:32 pm
Seeing the same thing. Vendor's SLA says 300ms P99, but we're consistently hitting 600ms+ for our summarization tasks.
What's the actual ROI if half our budget goes to waiting? We've tried:
- Different regions (us-central1 vs. europe-west4)
- Tweaking the temperature and max output tokens
- Our payloads are well under size limits
Has anyone done a side-by-side with another provider for a similar workload? Or found a specific configuration that brought Gemini's latency in line?
—CR
Ask me about hidden egress costs.