Skip to content
Notifications
Clear all

Where to start if I want to test its accuracy on highly technical documents?

11 Posts
11 Users
0 Reactions
21 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
Topic starter   [#21511]

Alright, let's cut through the usual "it's magical!" marketing fluff. We've all been burned by an AI confidently hallucinating API endpoints that don't exist or inventing configuration syntax that would bring down a production cluster. So, before I even consider letting something like NotebookLM near our internal architectural docs or post-mortems, I need a rigorous, cynical test plan.

My starting point is a corpus of documents where I already know the ground truth. I'm thinking of three tiers:
1. **Public, complex technical specs:** I'll feed it something like the AWS IAM JSON policy grammar reference or a Kubernetes API deprecation guide. I'll ask it to generate a valid policy or manifest based on specific, nuanced constraints from the doc. The failure mode here is subtle: will it correctly handle condition keys or API versions, or will it quietly make up plausible-sounding nonsense?
2. **Internal runbooks with known errors:** I have a few old, intentionally flawed Terraform modules (like a misconfigured `aws_subnet` resource) and their accompanying post-mortem analysis. The test is whether it can, when asked specific questions about the failure, correctly identify the root cause *by citing the exact lines and concepts from the post-mortem*, not by generating a generic "here's a common networking error" answer.
3. **Dense code/configuration combinations:** A documented CI/CD pipeline using GitHub Actions with specific security hardening steps. I'll ask it to modify the pipeline for a new use case while preserving all security controls. The key is to see if it understands the *interdependencies* spelled out in the docs, or if it just does a naive text splice that breaks the entire context.

The real metric isn't if it gets it right when the document is simple. It's how it fails when the document is complex. Does it admit uncertainty, or does it dress up a hallucination in convincing technical jargon? I want to see its confidence scoring and source anchoring under pressure.

I plan to run these tests next week. If anyone has already performed similar sanity checks on their own arcane infrastructure docs, I'm all ears. What were your failure modes? Did you find it was actually referencing the provided source, or was it quietly supplementing with its general (and often outdated/wrong) training data for our field?

-- cynical ops


Your k8s cluster is 40% idle.


   
Quote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Glad to see someone else who doesn't treat the release notes as gospel. Your first tier, public technical specs, is a solid start. But everyone tests against AWS and K8s docs, the vendor's own marketing team probably does that. The hallucinations get really creative when you move to lesser-known, poorly formatted RFCs or outdated vendor PDFs.

And on your second point about flawed internal runbooks, you're on the right track but maybe being too kind. The real test isn't if it can spot a known error in an old post-mortem, it's whether it will then propagate that error when you ask it to draft a *new* runbook based on the same flawed logic. I've seen tools perfectly diagnose a problem and then turn around and repeat it in the "solution".


cg


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Exactly. Testing on well-known docs is like a vendor's smoke test, it's table stakes. Your point about >poorly formatted RFCs or outdated vendor PDFs< hits home. That's where these tools choke on the noise, and where you actually need them to work.

I'd add that the cost of that hallucination changes with the source. An error in a public RFC might just break a prototype. But an error hallucinated from an internal, deprecated security guide could lead to a compliance finding. The fidelity of the document format matters less than the business impact of getting it wrong.

Has anyone tried feeding it scanned, fax-quality spec sheets from old hardware? That's my next brutal test.



   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

That's a really solid foundation for a test plan. I like that you're starting with a known ground truth - that's the only way you'll spot the subtle hallucinations.

Your second tier, testing against flawed internal runbooks, is particularly smart. It moves beyond "can it read" to "can it reason correctly." The real value in these tools isn't just retrieval, it's synthesis. If it can't spot the flawed logic in a post-mortem and then avoids repeating it in a new draft, it's dangerously overconfident.

One suggestion: maybe track not just if it makes an error, but *why*. Does it miss the nuance in a condition key because it didn't cross-reference a table? Does it gloss over the error in the runbook because the phrasing was too ambiguous? That diagnostic could be more useful long-term than a simple pass/fail.


Stay constructive


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You're right about the subtle failure mode. It's never the glaringly wrong syntax, it's the `"Effect": "Allow"` paired with a condition key that's valid in AWS IAM but from a totally different service. The plausibility is what wastes hours.

Your second tier is the real meat. The old, flawed Terraform module is perfect, but I'd go a step meaner: feed it the post-mortem *without* the module code itself. Ask it to generate a corrected module based *only* on the flawed narrative in the post-mortem. That's the synthesis test, and where most of these tools fall flat. They'll parrot the error from the analysis doc and bake it into new code.


Speed up your build


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Good point about the vendor testing. Makes their results meaningless for real use.

The real test is feeding it a mix of sources: a correct spec sheet plus a forum post with a popular but wrong workaround. Does it surface the forum post as a conflicting opinion, or does it just blend them into a franken-config that looks plausible? I've had tools cite a StackOverflow answer from 2015 for an Azure CLI flag that was deprecated in 2017.



   
ReplyQuote
(@brianw5)
Reputable Member
Joined: 3 months ago
Posts: 276
 

You're spot on about the subtle failure modes with condition keys and API versions. That's where the rubber meets the road. I'd take your K8s API deprecation guide test one step further and ask it to generate a migration path for a specific manifest across three minor versions, using *only* the deprecation notes. That's where I've seen even good tools trip up - they'll correctly flag a deprecated field but then suggest a replacement that was itself deprecated two versions ago, because they're just pattern-matching without chronological awareness.


Automate all the things.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That's a smart approach, starting with a known ground truth. I'd suggest a small adjustment to your first tier test, though.

When you ask it to generate a valid policy from the AWS IAM reference, don't just give it the spec. Give it the spec *and* a few conflicting, unofficial blog posts from 5+ years ago that are still highly ranked in search. The real test is whether it correctly prioritizes the primary source over the outdated, but plausible, community content.

Otherwise, you're just testing its reading comprehension, not its judgment. The synthesis part is where things get messy.


Keep it civil, keep it real.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your ground-truth approach is sound, but I'd argue your Tier 1 is still too clean. Feeding it a pristine AWS IAM spec tests parsing, not synthesis under pressure.

The real test is to create a "poisoned" corpus. Take that IAM reference, splice in a single, subtly incorrect sentence from a 7-year-old Stack Overflow answer (e.g., a subtly wrong condition key operator), and then ask it to generate a policy. Does it spot the contradiction and reject the bad source, or does it blend them and produce a policy that's syntactically valid but logically flawed? That's the microhallucination that costs a week.

Your second tier is the killer app. I'd add a specific latency metric: time from asking "what went wrong here?" to it correctly citing the exact line in the flawed Terraform module. If it takes 5 rounds of questioning, the latency to truth is too high for it to be a useful debugging partner.


--perf


   
ReplyQuote
(@isabell4)
Trusted Member
Joined: 2 months ago
Posts: 33
 

You've correctly identified the core failure mode with the >poisoned corpus< test. However, the effectiveness of that test depends heavily on the perceived authority of the poisoned source. A random Stack Overflow answer might be correctly discounted, but what if the incorrect statement is spliced from an outdated AWS partner blog post, or a deprecated example in a legacy version of the vendor's own documentation? The tool's weighting algorithm for source credibility is the real black box.

On your latency metric, I'd refine it. The critical measure isn't just the number of rounds, but whether the tool proactively flags contradictions during the initial ingestion or synthesis phase. If I have to play debugger and ask "what went wrong," it's already failed its primary role as an analytical partner. The useful tool surfaces the conflict between the IAM spec and the Stack Overflow snippet when it generates the policy, not when I'm reviewing the output.


PM by day, reviewer by night.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your proposed test on internal runbooks with known errors is the right direction, but I'd suggest a more precise failure mode to measure. The risk isn't just that it misidentifies the root cause, but that it generates a *causal chain* that is logically consistent with the flawed document's narrative yet factually incorrect based on the actual system state.

For example, if a flawed Terraform module had a VPC CIDR block conflict, the post-mortem might incorrectly blame a race condition in the provider. A sophisticated tool might correctly "identify" that non-existent race condition because the text supports it, rather than detecting the CIDR issue which requires cross-referencing the module's actual resource blocks. The test should be whether its answer contradicts the *code artifact* when the *text artifact* is wrong.



   
ReplyQuote