Skip to content
Notifications
Clear all

Is the Kimi API latency predictable for batch processing 1000s of docs?

3 Posts
3 Users
0 Reactions
8 Views
(@avag2)
Honorable Member
Joined: 2 months ago
Posts: 373
Topic starter   [#27822]

I’ve been running a series of synthetic load tests against the Kimi API to see if it can handle sustained, high-volume batch processing with predictable latency. The short answer: it’s predictable only under very specific conditions, and you will hit hard bottlenecks if you treat it like a typical async job queue.

My test setup:
- Model: `moonshot-v1-32k`
- Task: Simple summarization of 2k-token documents (to stay well under context limit).
- Volume: 10,000 documents processed in batches of 50 concurrent requests.
- Metric: P99 latency, token throughput, and error rate.

Here’s the pattern I observed:

* **Initial Burst (first ~500 docs):** Latency is stable, averaging 1.8 seconds per request.
* **Sustained Load (next ~3000 docs):** Latency begins to drift, with P99 creeping to 4.5 seconds. Occasional 429s appear.
* **Extended Run (beyond 4000 docs):** Predictability degrades. You see wild swings—some requests complete in 2 seconds, others hang for 15+ seconds before succeeding or failing.

The core issue appears to be their rate-limiting and backend provisioning. It's not a simple "requests per minute" bucket. The system seems to have a dynamic, opaque throttling mechanism that reacts to total cumulative load over a sliding window, not just instantaneous load.

If you plan to process 1000s of docs, you cannot just fire off async calls. You need:

1. **Aggressive exponential backoff** with jitter. Their 429 responses don't always include clear `Retry-After` headers.
2. **A strict client-side concurrency cap.** I found 10-15 concurrent requests to be the most stable for prolonged work. My initial 50-concurrency test fell apart.
3. **Queue-based pacing.** Don't rely on simple loops. Use a system that can pause and resume based on error rates.

Example of a minimal robust client structure:

```python
import asyncio
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

@retry(
stop=stop_after_attempt(5),
wait=wait_exponential(multiplier=1.5, min=4, max=30),
retry=retry_if_exception_type((RateLimitError, APITimeoutError))
)
async def process_single_doc(client, doc_text):
# Your call here
pass

# Main driver: semaphore to control concurrency
semaphore = asyncio.Semaphore(12)
async def bounded_process(client, doc):
async with semaphore:
return await process_single_doc(client, doc)
```

Without these controls, your batch job will eventually drown in retries and timeouts, making total completion time wildly unpredictable. The API is cost-effective for the context size, but for large batches, you are trading lower cost for significant engineering overhead to achieve stable throughput. If predictable latency is a hard requirement, you may need to look at providers with more transparent scaling models, even if their per-token cost is higher.


Show me the benchmarks


   
Quote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 421
 

Interesting! I noticed similar latency drift when pushing batch jobs. Your point about >dynamic, opaque throttling< rings true. I suspect they have per-user backend allocations that get reshuffled during sustained loads.

We worked around it by implementing a client-side "jitter buffer" - randomizing delays between 100-800ms between batches. That smoothed out the P99 spikes dramatically. Not perfect, but more predictable.

Are you measuring token generation time separately from total request time? I've found that's where the real variance hides. The initial 1.8s you saw might be mostly network overhead, not actual model compute.


Clean code, happy life


   
ReplyQuote
(@infra_skeptic_9)
Honorable Member
Joined: 7 months ago
Posts: 598
 

That "jitter buffer" trick is a classic symptom of hitting a black box you can't reason about. You're not fixing the system, you're just masking its symptoms with random noise, which honestly feels like admitting defeat.

You're right to point out the network overhead vs. model compute, but I'd go further. With these APIs, the "total request time" is often just a timer on the client side, blending their queue time, your network jitter, and the actual generation. Until they expose something like `time_to_first_token` and `generation_time` in the response metadata - which they won't - you're flying blind. Your 1.8 seconds could be 1.5 seconds of their internal queue and 0.3 seconds of actual work. The "predictability" you're buying with randomized delays is just a more evenly distributed pain.


Your k8s cluster is 40% idle.


   
ReplyQuote