Skip to content
How do I validate t...
 
Notifications
Clear all

How do I validate the sources an AI agent cites in its investigation summary? It keeps hallucinating CVE details.

1 Posts
1 Users
0 Reactions
20 Views
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
Topic starter   [#9436]

Hey folks, I've been deep-diving into building an AI agent for automated alert triage, and I've hit a consistent snag. The agent's investigation summaries are *mostly* good, but when it cites specific external sources like CVEs, exploit DB IDs, or even internal ticket numbers, there's a concerning rate of hallucination. It'll confidently reference a CVE-2024-XXXXX that doesn't exist or an "internal security bulletin #12345" we never published.

This undermines the whole point of having an AI SOC analyst. I can't just take its word for it before kicking off a response playbook.

My current setup is an agent using a GPT-4-class model with a RAG system over our internal docs and a tool-calling ability to "look up" external sources. The problem is, the tools don't always get called, or the model seems to "fill in the blanks" from its training data.

What strategies are you all using to force source validation? I'm thinking along these lines:

* **Pre-citation parsing:** Running a regex over the agent's output to extract any CVE-like patterns or ticket IDs, then programmatically verifying them.
```python
# Simple example of extracting and validating a CVE
import re
import requests

def extract_and_validate_cves(text):
cve_pattern = r'CVE-d{4}-d{4,7}'
found_cves = re.findall(cve_pattern, text)
valid_cves = []
for cve in found_cves:
# Quick check against a known API
url = f"https://cveawg.mitre.org/api/cve/{cve}"
resp = requests.get(url)
if resp.status_code == 200:
valid_cves.append(cve)
else:
log.warning(f"Agent cited non-existent CVE: {cve}")
return valid_cves
```
* **Stricter tool grounding:** Forcing the agent to provide a "source receipt" (like the exact query and snippet) for any external claim before it's allowed to include it in the final summary.
* **Post-generation fact-checking layer:** A separate, lightweight process that cross-references all cited entities against known-good databases before the report is finalized.

Has anyone implemented something like this? I'm particularly interested in how you balance validation with keeping the agent's response time low for SOC use cases. Are there any open-source tools or libraries starting to address this specific issue?

-- Weave


Prompt engineering is the new debugging


   
Quote