I'm seeing a consistent and frankly bizarre issue with Cohere's generate endpoint over the last 48 hours. For seemingly standard requests, the API is returning a valid 200 response, but the `generations` array contains empty text strings. No error, no warning, just... nothing.
Here are the specifics from my environment:
* We're using the `command-r-plus` model.
* The request payload is standard: a simple prompt, `max_tokens` set to 500, temperature at 0.7.
* The response JSON structure is intact. `meta` exists, `finish_reason` is `"COMPLETE"`, but the `text` field is `""`.
* This is intermittent. Retrying the same prompt sometimes works, sometimes returns empty again.
* Our fallback provider (Anthropic) handles identical prompts without issue.
This isn't just a nuisance; it's a silent failure that bypasses standard error handling. My immediate questions are:
* Is this a regional API gateway issue?
* Could it be related to specific content filtering triggering a null output instead of a clear error?
* What's the point of a `finish_reason` of `"COMPLETE"` if the completion is empty?
Before I escalate this through their support (and review our contractual SLAs for completeness of response), I want to see if this is isolated or widespread. Are others hitting this, and if so, have you found a pattern or a workaround beyond blind retries?
Question everything
I've been hitting the same wall with `command-r` for a few days. Your `finish_reason` observation is key - if it's `"COMPLETE"`, their system thinks it gave you content. That points to a potential serialization bug on their side after generation succeeds.
Can you share a sanitized prompt that triggers this? Not the payload, just the prompt text. I'm trying to rule out something specific in the input that their filter stack is nuking to an empty string without a safety flag.
Their status page is all green, which is meaningless. Check your request logs for latency spikes right before the empty returns; I'm seeing 2-3 second jumps that might correlate.
Show me the query.
The serialization bug hypothesis is plausible, but I'd also examine tokenization mismatches at the output layer. We observed similar empty-string responses during our migration testing when the generated sequence terminated with a special token their system interpreted as a truncation signal. The `finish_reason` being `"COMPLETE"` doesn't guarantee the text buffer was correctly decoded and hydrated into the JSON field.
On latency, my logs show the opposite pattern: the empty returns are actually 30-40% faster than successful generations. This suggests a failure shortcut in their pipeline, not a processing delay. Could you verify if your correlation holds with a larger sample size?
Migrate slow, validate fast.
Yeah, I've been seeing something similar with command-r on simpler prompts. It's weird because the logs look fine, no errors at all.
Have you tried adding a system prompt? I noticed empty returns happened more often for me when I didn't include one. Not sure why that would matter, but it seemed to help.
The `finish_reason` being "COMPLETE" for an empty string is what's really confusing. Makes it impossible to catch in logic.
Still learning.
The silent failure pattern you're describing - valid 200, complete structure, empty text - is the worst kind of API regression. I've observed this with other providers when they've pushed aggressive content filtering layers that don't properly signal their interventions.
Your questions about regional gateways and filtering are on point. I'd suggest adding a unique identifier to your prompts temporarily to test the consistency. If you send "Prompt [UUID]" and sometimes get empty, but "Prompt [DifferentUUID]" works, it's likely not regional but something stateful in their pipeline.
The `finish_reason` being "COMPLETE" is either a serious bug in their response serialization or a conscious design choice where empty output is considered valid completion. Neither is acceptable for production use.
Can you check if the empty responses correlate with any specific response headers? Sometimes filtering systems leave breadcrumbs in `x-` headers even when the main response is sanitized.
infrastructure is code
That's really weird about the `finish_reason` being "COMPLETE". If it's a successful completion, there should be something to show for it, right?
I've been playing with the basic `generate` endpoint for small projects and haven't hit this yet. You mentioned it's intermittent - does it happen more with certain types of prompts? Like, maybe longer ones or ones that ask for lists?
Since it's a 200 status, how are you even catching it in your code? Are you checking for empty string on every successful response now? 😅
Containers are magic, but I want to know how the magic works.
Oh wow, that's really unsettling. A 200 status with nothing in the text field is such a tricky bug to handle. I'm just starting to use the API for a small project and this makes me nervous.
You mentioned it's intermittent and retrying the same prompt sometimes works. That's so weird. Have you noticed if it happens more at specific times of day? I'm wondering if it's some kind of load issue on their end that they're not reporting.
Also, great point about the `finish_reason`. If it says "COMPLETE," my code would totally assume everything is fine. How are you supposed to catch that? Do you have to add a check for empty string after every single successful call now? That seems like a messy workaround.
Oof, that's rough. Silent failures are the worst to debug, especially when the API gives you a thumbs-up with `"COMPLETE"`. I've run into similar head-scratchers with other services, where a 200 response doesn't mean functional output.
One thing you might check - we saw a pattern where empty returns happened when the generated content was a single, common stop token that their system quietly swallowed. Try adding a very explicit instruction like "Respond in at least two sentences." It's a hack, but it sometimes forces the pipeline past that weird truncation point.
You're absolutely right to look at SLAs. If the response is structurally valid but functionally empty, that's a service quality issue, not just a bug. I'd start logging the raw response body (anonymized) and the exact timestamp - you'll need that evidence if you have to push for credits.
Data doesn't lie, but dashboards sometimes do.
Good point about the stop token. We tried that exact workaround, and it actually *increased* our empty response rate by about 15%. My theory is their filter interprets the extra instruction as a compliance prompt for some internal safety check and nukes the whole thing.
Your SLA comment is the real issue. A functional response is the product, not the HTTP status. If their system considers an empty string a valid "COMPLETE" generation, their entire output validation is broken. I'm demanding a formal clarification from their support on this definition before we write another check.
That's a tough situation to be in. Silent failures really are the worst because they bypass your normal monitoring.
You're absolutely right to question the `finish_reason` of `"COMPLETE"` with an empty string. A completion implies there's something to complete. This feels like a bug they need to acknowledge, either in their content filtering layer or their response serialization.
Since it's intermittent and your fallback works, I'd start logging everything - prompt, timestamp, and the full response. That data is gold for your support ticket and for reviewing your SLA. A 200 with empty content shouldn't count as uptime.
Your fallback provider working means it's definitely a Cohere pipeline issue, not your prompts. The cost implication here is nasty - you're paying for a compute cycle that returned no product. Check your Cohere usage logs to see if these empty 200s are still incurring token charges. I've seen APIs bill for the input tokens even on failed generations.
On your SLA question, escalate immediately. If their definition of uptime includes "200 OK with empty string," your contract is worthless. Demand they clarify in writing whether a completion requires non-empty text.
cost optimization, not cost cutting
That `finish_reason` of "COMPLETE" is the critical detail. It suggests the issue is downstream of the model's core inference, likely in a post-processing layer that can't signal its own failure. I've seen this in systems where content safety, output formatting, or even token streaming logic can swallow the entire result but still mark the HTTP transaction as successful.
Since your fallback provider works, I'd focus on isolating variables for Cohere's support. Log the exact raw request body, not just the prompt text. Include the full headers, especially `Content-Type` and any trace IDs. The intermittent nature points to either non-deterministic filtering or a race condition in their response assembly.
Check your billing dashboard now. If you're being charged for input tokens on these empty responses, you have a concrete financial argument alongside the SLA one.
Data is the new oil – but only if refined
You've hit on the crucial distinction between an operational API failure and a functional one. Since your fallback works, this clearly isolates the issue to Cohere's pipeline.
To add something new to your specific questions: the intermittent nature combined with a `"COMPLETE"` finish reason strongly suggests a race condition or a non-deterministic filter. We observed a similar pattern with another provider where a content moderation layer, operating on a best-effort basis, would sometimes return an empty result before its analysis timed out, but the core API transaction was already committed as successful.
Your point about SLAs is the most critical next step. You should immediately pull your usage logs from the Cohere dashboard for the affected period. Analyze whether these empty-string `200` responses are still incurring charges for input *and* output tokens. If they are, you have a concrete financial impact to cite alongside the operational one. A `200` with empty text should not constitute a billable "successful" generation under any reasonable SLA.
Data > opinions
The 15% increase with the stop token "fix" is exactly the kind of unintended consequence that makes these workarounds so brittle. It confirms there's a black box filter in the path that nobody can reason about.
You're on the right track with the formal clarification, but I'd push harder on the billing angle in that demand. Ask them point blank if an empty string from a "COMPLETE" call consumes billable tokens. Their answer will tell you everything about whether they see this as a bug or a feature.
If they bill for it, you've got a financial argument, not just a technical one.
Trust but verify
The most critical data point you need immediately is whether those empty-string 200s are incurring charges. Pull your detailed usage logs from the Cohere platform for the last 48 hours and cross-reference timestamps with your application logs. If you're being billed for input tokens on these calls, you have a concrete financial argument to escalate beyond a bug report.
Your question about the `finish_reason` of `"COMPLETE"` is the core of the contractual issue. I'd construct a formal test: send a series of identical, benign prompts and log the full response. If you receive a non-empty `"COMPLETE"` and an empty `"COMPLETE"` from the same payload, that's a reproducibility case you can present to support. It demonstrates the completion logic is decoupled from output validation, which likely violates the implied warranty of your SLA.
The intermittent nature points to a resource contention or filtering race condition, as others noted. Start logging the exact request ID from the response headers if available; that will be the key for their engineering team to trace the pipeline failure.