Just trialed the new AI agent on our support desk. It's defaulting to Spanish responses for about 30% of tickets, despite all system settings and historical tickets being in English.
This isn't a minor bug. It creates a compliance headache for us in regulated markets. The vendor's response was to check our language settings, which we already did. Now they're suggesting a "custom training" add-on to fix it. That's a hidden cost on a broken feature. Anyone seeing similar language drift with their implementation?
Your observation about language drift isn't isolated. I've documented similar behavior in multilingual benchmarks where the model's context window includes user queries in Language B, even when the system prompt explicitly mandates Language A for output. The issue is often data contamination in the training set or a poorly implemented token bias at inference time.
The vendor's suggestion of a custom training add-on is a red flag. It indicates the base model has a fundamental localization flaw they can't patch. You shouldn't pay for that. A properly architected system should respect the language instruction in the prompt without additional fine-tuning. I'd push back and demand they disclose the exact model version and any regional token weighting they're applying.
Have you logged the specific user tickets that triggered Spanish responses? There might be a pattern, like certain technical terms or names that share spelling with Spanish words, causing the model to latch onto a wrong language context.
numbers don't lie
The compliance angle is the real kicker. That moves it from "annoying bug" to "can't go live." Seen it happen when the vendor's backend pools inference across clients to save costs. Your English tickets might be hitting a GPU cluster trained for another customer's Spanish data.
Prove it.
"Hidden cost on a broken feature" is exactly right. We see this pattern with monitoring tools too, where vendors try to monetize fixing a core metric.
Your compliance risk is real. If you can't trust the system prompt, you can't audit it. I'd halt the trial now and demand they treat it as a critical bug, not a paid add-on.
Check your raw request logs. If the language header is set correctly and it's still happening, their inference layer is defective.
null
Right? That "monetize the fix" pattern is everywhere now. We saw it with a major CRM's duplicate merge logic. The base feature had a known flaw that created orphaned records, and the official solution was a premium "data integrity" module.
Your point about the logs is key. If the header's correct, the defect is in their runtime. I'd even sniff the actual API call with a proxy to confirm they aren't stripping the instruction. It's frustrating when you have to prove their system is broken.
That proxy idea is good, but it's just more unpaid work for you. If you're at the point of inspecting raw packets to verify their API, you've already lost. The contract should state the output adheres to the input language header. If it doesn't, it's a breach. Sniffing traffic is for debugging your own junk, not a paid vendor's.
If it ain't broke, don't 'upgrade' it.