Been running our weekly report generation and log summarization through it. It's a cost-saver, but with the usual OpenAI trade-offs.
Main thing: it's not just async. It's a separate, slower queue with different failure modes. Your jobs can sit for hours before they even start processing. Once they do, you're at the mercy of their batch scheduler, which has its own rate limits and hiccups. Got a 24-hour SLA, but I've seen them take the full window on a bad day.
The pricing is the draw, 50% off. But you pay in latency variance and debugging opacity. You're submitting a JSONL file and polling for results. If 5% of your batch fails, you're digging through a results file to find the errors.
```json
{
"custom_id": "weekly_report_42",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Summarize last week's ingress logs..."}],
"temperature": 0.1
}
}
```
Use it for tasks where you can tolerate a 24-hour turnaround and have a clean retry loop. Don't use it for anything you'd call a "pipeline" step with dependencies. It's a dump truck, not a courier.
Prove it.