Hi everyone. New to the forum but been testing Hailuo for a project management doc review.
Our team just processed 1,000 project briefs and status reports through Hailuo's extraction API. We were really counting on the advertised 99% accuracy for dates, owners, and key deliverables.
Our manual audit of a 300-document sample shows our actual accuracy is around 87%. That's a big gap, especially for missed dependencies.
Has anyone else done large-scale accuracy tests? Curious if we set something up wrong or if others see this too. We used the default settings for document processing.
Thanks in advance!
Still learning.
Thanks for sharing those numbers, that's quite a gap. Our team saw something similar on a smaller batch of Jira export summaries last month, though our variance was closer to 91%. I'm curious if the discrepancy is more pronounced with certain document types.
Has anyone in the community run a comparable test with a tool like Docparser or Parseur for similar project documentation? I'm trying to understand if this is a Hailuo-specific tuning issue or a more general challenge with extracting dependencies from informal status reports.
A 1,000-document sample is a solid test volume, and that 87% figure is a significant datapoint. I've found vendor accuracy claims often stem from highly sanitized internal benchmarks, usually on clean, formatted documents that don't reflect real-world project chaos.
Your mention of dependencies is key. In my own benchmarking, extraction accuracy for simple fields like dates and names is often high, but relational data like dependencies, where context is scattered across paragraphs or bullet points, sees a notable drop. The advertised 99% likely refers to a controlled extraction task, not a composite score across all field types you're using. Did you break down your accuracy by field category? I'd expect the "key deliverables" and "dependencies" to be dragging the average down considerably.
Using default settings is also a potential factor. For complex project docs, you usually need to provide explicit examples or tweak the confidence thresholds. Without that, the model makes more guesses on ambiguous phrasing, which directly impacts reproducibility.
-- bb42