The legal data licensing point that user1085 and others brought up is crucial, and it makes me think your evaluation needs a new layer. You're looking for error typology, but those errors might be traceable to specific *types* of licensed content. If a vendor's data is heavy on legal transcripts but light on technical manuals, you'd see a particular pattern of pragmatic errors in procedural scripts.
For the side-by-side comparisons, I haven't seen published benchmarks for video tools specifically. But based on my experience localizing ERP system training materials, the drop-off user364 mentioned for non-English pairs is severe. We tried a quick test of a German to Portuguese video for a warehouse process, and the tool substituted a general term for a specific bin location code. That's a critical omission in our context.
Your methodology question is the tough part. The bilingual speaker scoring on intent seems standard, but how do you control for the reviewer's own domain knowledge? A translator might score something as semantically correct, but if they don't know supply chain logistics, they might miss a subtle but catastrophic shift in meaning for an operational instruction.
That's a really smart way to break down the errors. In your tests, did you also see a lot of syntax problems? I'm wondering if that's more common in distant pairs.
For evaluation, maybe you could score each error type separately? Like, a script might have perfect word choice but totally fail on nuance. That breakdown could show where a tool is actually failing.
Still learning
Syntax errors are absolutely more pronounced in distant pairs, but I've observed they manifest differently than simple grammar mistakes. The problem is often at the clause or sentence structure level, where the model imposes the source language's syntactic order onto the target language, creating technically "correct" but highly unnatural phrasing. In a Japanese to English translation of a technical document, we saw repeated instances of the verb being placed at the end of a long conditional clause, a direct carryover from Japanese SOV structure that made the instructions confusing to parse quickly.
Scoring error types separately is the only way to get a useful signal. In our internal bake-off, we used a weighted scoring matrix across five categories: lexical accuracy, syntactic fluency, semantic fidelity, pragmatic appropriateness, and formatting integrity. A tool could score 90% on lexical but 40% on pragmatic, which tells you it's fine for basic UI strings but dangerous for customer facing content. The nuance failures user169 mentioned are almost entirely captured in the pragmatic category.
The real challenge is that syntax and pragmatics are often intertwined. An awkward syntactic structure can change the perceived tone or urgency of a procedural step, which then bleeds into a pragmatic error. Isolating them requires very careful rubric design.
No free lunch in cloud.
Your weighted scoring matrix is a solid approach, and the interplay between syntax and pragmatics is exactly where the most revealing benchmark data hides. We ran a similar test suite on a set of financial report translations, and the results underscore your point.
A tool scoring highly on lexical and syntactic metrics for English-to-German translations would still produce pragmatically disastrous output because it couldn't handle the shift in formality register required for different sections. The syntax was flawless, but the choice of address (formal vs. informal "you") was randomly applied, which in that context completely alters the authority of the document.
This suggests any useful benchmark must use source material with *intentional* pragmatic complexity, not just technical complexity. Generic news text won't surface these failures.
-- bb42
Of course accuracy isn't uniform. It's a data problem, not a linguistics one. Your "semantic fidelity" test is chasing the wrong metric.
Benchmarks based on linguistic distance are academic. The real map is what content was licensed. If a vendor's training data is 90% movie subtitles and 10% technical docs, your procedural script will fail regardless of language pair. English to Japanese might outperform English to Spanish if the vendor bought a better J-drama corpus.
Everyone's looking for errors in the model. Look at the vendor's data sources first. The rest is just symptoms.