Just learned something cool for ChatPDF. I'm always feeding it AWS whitepapers or internal architecture docs, and it sometimes gets confused by all the acronyms (VPC, IAM, ECS, etc.).
Someone in my team suggested uploading a simple glossary file first—like a one-pager with "VPC = Virtual Private Cloud"—before the main document. Tried it with a GCP pricing guide and it worked way better! The answers were more accurate because it understood the terms from the start. Has anyone else tried this trick? Seems like a simple way to improve the context it has.
Still learning
That's a solid tip. I do something similar with our internal knowledge base articles. Instead of a separate file, I sometimes paste the glossary at the top of the main document before uploading. It seems to anchor the context.
One caveat: if you're using a system with a limited context window, that glossary eats into your token budget. For really long docs, you might have to prioritize which terms are most critical. But for a one-pager like you described, it's perfect.
Automate the boring stuff.
Oh that's a smart trick! I've definitely had similar issues feeding Terraform module docs to these tools. They'd get tripped up on HCL-specific terms.
I wonder if it helps more with acronyms or with proprietary product names. Like, would it better recognize "AWS Transit Gateway" versus just "Transit Gateway" in a doc? Might have to test that.
Great tip for the team's internal runbooks, too. Thanks for sharing!
Infrastructure as code is the only way
This is a useful technique, especially for the type of dense infrastructure documentation you're working with. The effectiveness likely stems from providing a fixed reference point, which is crucial for a model parsing ambiguous acronyms like AZ (Availability Zone versus Amazon's other uses).
The method mirrors a common troubleshooting step in configuring entity recognition for older document processing pipelines, where you'd prime the system with a known dictionary. A practical caveat for longer sessions is that the model's attention can drift over many interactions, so for a complex analysis, you might need to tactically re-insert the key term definitions in a follow-up prompt rather than relying solely on that initial glossary upload.
I've found this is particularly critical when dealing with vendor-specific terms that have generic homonyms, like "Gateway" or "Cluster." Defining them upfront reduces the chance of the model applying a generic computing definition to, say, an AWS Direct Gateway.
Plan the exit before entry.