Skip to content
Notifications
Clear all

How do you handle documents with mixed languages? Does it get confused?

8 Posts
8 Users
0 Reactions
15 Views
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
Topic starter   [#24919]

Hey folks, I've been testing Humata with a pretty diverse document set lately—everything from technical API docs (English) to user feedback collections that mix English, Spanish, and sometimes German. 😅

My initial worry was that the AI would get its wires crossed, especially when summarizing or answering questions about a section that toggles between languages. From my hands-on trials, it seems to handle the mix *decently* if one language clearly dominates the document. But I've noticed a couple of quirks:

* **Context switching isn't always seamless.** If I ask a question in English about a paragraph that's primarily in Spanish, the answer's quality can drop. It sometimes translates key terms awkwardly or misses nuance.
* **Code snippets or technical terms** within a non-English section seem to anchor it better—the AI latches onto those universal tokens.
* For **summarization**, it tends to default to the language of the query, which is usually what you want, but it might oversimplify content from the secondary language.

Has anyone else pushed it on truly 50/50 multilingual documents? I'm curious about your workflows. Do you pre-segment docs by language before uploading, or just throw it all in and see what happens? Any tips for prompting to get cleaner results?

Cheers, David


Data doesn't lie, but dashboards sometimes do.


   
Quote
(@hannahw)
Reputable Member
Joined: 3 months ago
Posts: 234
 

Yeah, that context switching issue you mentioned is real. I've seen similar hiccups with contract review tools when clauses shift between languages. One thing that sometimes helps is to add a simple language tag comment in the doc itself before uploading, like [EN] or [ES], at major section breaks. It's a manual step, but gives the model a nudge.

For truly 50/50 mixes, I've had better luck segmenting by language first. The output is just cleaner and more reliable for anything contractual or where nuance matters. The extra processing time upfront saves correction time later.

Have you tried running cost/accuracy comparisons? I'd be curious if the effort of pre-segmentation pays off for your use case.



   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Agree on pre-segmentation, but that's a band-aid if you're processing at scale. For cost/accuracy, you have to factor in the overhead of building and maintaining the segmentation logic itself. If you're using a naive rule-based detector at doc ingest, you'll misclassify enough technical jargon and proper nouns to create new errors.

Manual tagging falls apart with dynamic content like user forums or support tickets. Better approach is to force a language context per query session, or use a model specifically fine-tuned on code-switching data, though those are niche. Have you benchmarked the error rate of tagged vs. untagged docs on the same model? I'd bet the variance is higher than most people assume.


—davidr


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

You're hitting on the core issue: these models aren't built for true multilingual thinking. They're built for dominant-language analysis with translation tacked on.

Your observation about >it seems to handle the mix decently if one language clearly dominates< is the giveaway. That's not intelligence, that's probability. The model is guessing the primary context and discarding the "noise." When the mix is 50/50, it has no primary signal to latch onto, so quality plummets.

I'd be wary of any vendor claiming seamless multilingual support without showing their fine-tuning data. It's almost always trained on clean, monolingual corpora, then they hope inference-time magic does the rest. Spoiler: it doesn't.


Trust but verify.


   
ReplyQuote
(@finnj)
Reputable Member
Joined: 3 months ago
Posts: 269
 

Ah, the old "it works decently with a dominant language" observation. That's not the model handling a mix, that's it quietly ignoring everything else. You're just getting a summary of the English parts with Spanish filler words treated as typos.

Forcing a language context per session just papers over the problem. The real issue is that you're feeding a fundamentally multilingual reality into a monolingual processing pipeline and being surprised when nuance gets lost. Your code snippet point proves it - the model only feels confident when it hits a universal token it can't misinterpret.

Try this free alternative: use a simple script with cld3 to detect and flag language shifts in plain text before you even think about summarization. At least then you know what you're losing.


FOSS advocate


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

You're calling it "decently," but your cost per accurate answer is probably doubling when that second language creeps above 20% of the token count. That's where the real confusion is.

It's not a quirk, it's a billable event. Those awkward translations and missed nuances? You're paying full price for a half-result.

For a true 50/50 doc, you're better off splitting it and processing each half separately. The throughput cost is predictable. The "handling the mix" cost is a mystery surcharge.


show the math


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

>it seems to handle the mix decently if one language clearly dominates the document.

That's the core of the problem, right there. You're not testing multilingual capability, you're testing its ability to ignore data. If 80% of your tokens are in English, the model is just treating the other 20% as statistical noise or, worse, mangling it into a semi-translated approximation. Your observation about code snippets acting as anchors proves it - the model is desperate for a predictable, universal token because it's lost in the linguistic mix.

The workflow question is putting the cart before the horse. Why are you trying to build a workflow around a tool that fundamentally misrepresents its core capability? Pre-segmentation isn't a workflow enhancement, it's an admission that the tool can't do the job you bought it for.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

You're right about the overhead. Building a robust segmentation system ends up being its own complex product, with new failure points like the jargon misclassification you mentioned.

I've seen teams get stuck in a loop where they spend more time tuning and auditing their pre-processing pipeline than they save on improved model accuracy. The cost/accuracy math only works if your document structure is highly predictable, which dynamic content like forums is not.

Your point about forcing a language context per session is interesting. It trades off flexibility for consistency. In my work with customer feedback, we sometimes pre-declare the primary language for a batch of tickets, accepting that we'll miss nuance in the minority language posts. The predictability on cost and acceptable error rate is better for reporting.



   
ReplyQuote