I've been conducting an in-depth evaluation of Humata.ai for potential enterprise adoption, with a specific focus on its document interrogation and citation accuracy. A recurring and significant issue I've identified in my testing is the inconsistency of cited page numbers. The problem isn't that the citations are entirely wrong, but that they are frequently off by one or two pages, which in a legal or compliance context renders the feature nearly unusable.
My workflow involves uploading complex technical whitepapers, security audit reports, and lengthy SaaS contracts—all PDFs with clear pagination. When I ask a question like, "What are the data retention obligations in this agreement?" Humata will often provide a correct answer and cite, for example, "page 12." However, upon manual verification, the relevant clause is actually on page 11 or 13. This error margin suggests a systemic problem with the document parsing or page anchoring logic.
I have attempted to isolate the variables:
* **Source Document Format:** The issue persists across PDFs generated from Word, scanned documents with OCR, and even simple text-based PDFs.
* **Question Complexity:** It occurs with both simple fact retrieval ("What is the SLA percentage?") and more complex synthesis ("Summarize the termination for cause clauses").
* **Document Length:** Shorter documents (sub-20 pages) seem slightly less prone, but the error is not eliminated.
This isn't a minor formatting quirk. For any professional use case—particularly in vendor evaluation and contract negotiation where I operate—the integrity of the citation is paramount. A page number off by one can mean referencing a definition instead of the obligation, or a warranty disclaimer instead of the warranty itself.
I am seeking to understand:
* Is this a known issue with a documented workaround?
* Has anyone reverse-engineered the conditions that lead to stable, accurate citations? For instance:
* Does pre-processing the PDF (e.g., ensuring every page has a visible numbered footer) improve accuracy?
* Is there a correlation with document sections that have cover pages or large images?
* Does the "chunking" or segmentation method Humata uses explain the off-by-one error?
* Are there any official statements or configuration parameters that address citation fidelity?
My current assessment is that while the NLP engine is powerful, the citation inaccuracy presents a high TCO risk due to the required manual verification step, which negates much of the efficiency gain. I'm interested in any rigorous analysis or reproducible findings from other users who rely on precise sourcing.
—LJ
—LJ