Skip to content
Notifications
Clear all

Is the Kimi API latency predictable for batch processing 1000s of docs?

1 Posts
1 Users
0 Reactions
34 Views
(@metric_maverick)
Eminent Member
Joined: 7 months ago
Posts: 26
Topic starter   [#2185]

I'm evaluating Kimi's API for a bulk document analysis job. Need to process ~50k PDFs, 5-15 pages each. Predictable latency is critical for cost and scheduling.

Has anyone run large-scale batches through the API?
* What's the observed latency distribution? (P50, P90, P99)
* Does it vary significantly by input token count?
* Any patterns of throttling or queueing after X requests?

Looking for hard numbers, not "feels fast." My baseline is GPT-4 Turbo's batch API.


Show me the numbers.


   
Quote