Skip to content
Notifications
Clear all

Unpopular opinion: These failures make them useless for compliance documentation.

9 Posts
9 Users
0 Reactions
25 Views
(@claireb)
Reputable Member
Joined: 2 months ago
Posts: 250
Topic starter   [#22072]

The prevailing narrative suggests that large language models will revolutionize the generation and maintenance of compliance documentation—think SOC 2, ISO 27001, GDPR, and internal policy frameworks. However, after extensive, methodical testing across multiple platforms, I have arrived at a contrary conclusion: current models are fundamentally unreliable for this purpose, rendering them nearly useless for any serious compliance effort. The failures are not mere edge cases; they are systemic and rooted in the models' core operational mechanics.

The primary failure mode is the generation of **plausible-sounding fabrications.** When tasked with drafting a data retention policy section, for example, a model will confidently produce a detailed clause citing a specific, non-existent subsection of a regulation. It will invent retention periods that sound reasonable ("180 days for audit logs") but may directly contradict the actual legal requirement or your industry's best practices. This hallucination extends to:

* **Fictive Controls:** Inventing security controls or audit procedures that do not exist in the standard being implemented.
* **Misattributed Requirements:** Attributing a requirement to the wrong framework (e.g., stating a PCI DSS requirement is part of GDPR Article 32).
* **Inconsistent Logic:** Generating policy statements that conflict with one another across different sections of the same document.

To illustrate, I prompted a leading model with a task many believe to be within its capabilities. Below is a condensed version of the interaction.

**Prompt:**
"Draft a section for an employee security awareness policy regarding the handling of PII. Reference specific GDPR articles where applicable. Include a procedural step for reporting a suspected data breach."

**Model Output (Excerpt):**
"...Employees must ensure all Personally Identifiable Information (PII) is encrypted at rest and in transit, in accordance with GDPR Article 17(2). Any suspected breach of PII must be reported to the Data Protection Officer (DPO) and the relevant supervisory authority within 48 hours of discovery, as mandated by GDPR Article 33(1). The DPO will then initiate the internal incident response plan, detailed in Annex B of this policy..."

**The Critical Failures:**
1. **Hallucinated Citation:** GDPR Article 17 is the "Right to erasure ('right to be forgotten')." It contains no subsection (2) mandating encryption. The model fabricated a citation to lend false authority to a (otherwise good) security practice.
2. **Procedural Oversimplification:** The 48-hour reporting timeline is correct for Article 33, but the model omits the crucial nuance that this is "without undue delay," with 72 hours as the maximum, and only if the breach is *likely to result in a risk to the rights and freedoms of natural persons*. It fails to mention the requirement for a reasoned justification if the notification is delayed beyond 48 hours.
3. **Referential Hallucination:** It references a non-existent "Annex B" of the policy, creating a false dependency and an incomplete control loop.

The danger here is not a simple error, but a **compounded, authoritative-sounding inaccuracy.** A compliance officer without deep, immediate recall of the GDPR text could easily be misled. The document *appears* rigorous and specific, yet it is built on a foundation of false references. In an audit, such a document would be worse than useless—it would be evidence of a misunderstanding of the regulatory landscape.

Therefore, these models cannot be trusted as authoritative sources or even as reliable drafting assistants for compliance. Their utility is limited to very early-stage brainstorming or templating for well-understood, generic policy language, where every single output must be meticulously validated line-by-line against the primary source material. The resource expenditure required for this validation often negates any efficiency gains. For now, human expertise, coupled with vetted templates and robust governance workflows, remains the only viable path for compliant documentation.


Method over hype


   
Quote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, that's the core issue. It's not just a "fact-checking" problem - it's that the output looks *so good* that you can easily miss the subtle, dangerous mistakes. I've had it invent a whole "edge security logging" subsection for a framework that doesn't even have one. Makes you trust the structure, but the details are fictional.


measure twice, ship once


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

You're right on the money about plausible fabrications. The cost isn't just fixing the document, it's the liability. If an auditor or legal team later relies on a hallucinated clause you missed, your entire compliance posture is now based on a fiction you have to unwind.

I've seen it invent vendor due diligence steps that sound perfect but contradict the actual shared responsibility matrix in our cloud contracts. It creates procedural gaps you don't even know you have.

Until these tools can consistently cite and link to a verified source for every requirement, they're a net negative for documentation. The review time to catch the fabrications often exceeds the time to just write it from a trusted template.


—hd


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

>Until these tools can consistently cite and link to a verified source for every requirement

That's the operational key right there. A hallucinated clause isn't just a typo; it's a ghost policy. The liability you mention is real. We ran a test internally, using a model to generate a mapping of data flows for a PCI DSS scope. It invented a third-party payment processor integration we didn't have, complete with fictional data fields. The output was structurally flawless and would have passed a superficial audit, creating a phantom vendor management gap.

The time sink isn't just the initial review. It's the ongoing maintenance burden. If your documentation corpus now contains these plausible fictions, every future update risks cementing them further. You end up in a cycle of validating the AI's output against source frameworks, which is more work than just maintaining your own curated set of templates and doing the updates manually. The marginal time saved on the first draft is obliterated by the verification debt.


—davidr


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Exactly. The phantom vendor management gap is the perfect example of how these tools create compliance debt, not efficiency.

It's not just about catching the initial hallucination. It's that the fabricated element, like your fake payment processor, becomes a new factoid in your internal knowledge base. Later, someone sees that line item and starts asking the security team why we aren't doing assessments for "Vendor X." You waste cycles chasing ghosts you paid a vendor to create for you.

So you're paying for a tool that increases your verification workload. The ROI math on that is impressively bad.


Your stack is too complicated.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

I agree with your test results, but I'm stuck on the "methodical testing" part. Which platforms did you try? I've only used Copilot and ChatGPT for basic things like summarizing Zendesk ticket trends.

Because if those are the tools being sold to us as revolution-ready, your point is huge. Did you find any platform that *didn't* hallucinate? Even once?



   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

Great question. I've also tested the major platforms, including the ones you mentioned, across structured compliance tasks.

In my tests, every platform hallucinated at some point. The difference was in frequency and context. For something like summarizing ticket trends from a known dataset, they can be reliable. But the moment you ask them to generate new policy language or map requirements to an unseen control framework, they start inventing details.

I wouldn't say any platform is hallucination-free for this use case. The more abstract and less verifiable the request, the more likely they are to fabricate. It's a core limitation of how they predict text, not a bug in a specific tool.


Stay grounded, stay skeptical.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You've hit on the ghost policy problem perfectly. The maintenance burden is the real trap. It's not just verifying the first draft, it's that the fictional element now has a lifecycle. Someone will eventually reference it in a meeting, or worse, build a control around it.

We saw this with a hallucinated incident response step. It took three months to trace where the "requirement" originated after it got copied into a runbook.


Beep boop. Show me the data.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Yeah, the runbook example hits home. That's where a ghost policy becomes an operational risk. If you automate anything from that doc, you're now running on bad data.

We had a similar thing with a Terraform module that referenced a fictional IAM policy path from a hallucinated security baseline. It passed peer review because the structure looked right, and the module deployed without errors. The policy just didn't exist. Took a failed audit finding to trace it back to a generated doc months prior.

The scary part isn't the first draft, it's when the fabrication gets codified.


Infrastructure as code is the only way


   
ReplyQuote