Everyone's suddenly worried about "hallucinated" citations, but most of the evaluation frameworks I see are just checking for a URL format regex. That's useless. A model can hallucinate a perfectly formatted, plausible-looking URL that points to nothing or to a completely unrelated page.
You need to test for *actionable* fabrication. Here's my method.
First, generate a set of prompts that explicitly demand citations on a diverse set of topics (e.g., "Give me sources for the impact of vector databases on LLM latency"). Use a known, controlled dataset of real documents for your test—maybe a snapshot of Wikipedia or a specific set of research papers. This gives you ground truth.
Then, run your prompts and extract every URL or citation. The evaluation has two layers:
* **Syntactic Check:** Does the output look like a citation? This catches lazy hallucinations.
```python
# Simple regex is just the first filter
import re
url_pattern = re.compile(r'https?://(?:[-w.]|(?:%[da-fA-F]{2}))+')
# But this is table stakes.
```
* **Operational Check:** This is the critical part. For each extracted citation:
* Verify the domain/host exists (DNS lookup).
* Perform a HEAD/GET request (with rate limiting and respect for robots.txt) to check if the resource actually returns a successful status code (e.g., 2xx).
* For non-URL citations (e.g., "Smith et al., 2023"), check if it can be matched to an entry in your controlled ground-truth dataset.
The metric is the **hallucination rate**: (Number of fabricated or unresolvable citations) / (Total citations generated). You must also track partial hallucinations, where a real base URL is given but with a fake path or query.
Most vendor benchmarks skip the operational check because it's slow and requires networking. That's why they're misleading. If you're not actually trying to fetch the thing, you're not evaluating the real risk.
-- bb
-- bb
Good operational start, but verifying DNS is the cheap part. The expensive part - both computationally and for your cloud bill - is the actual GET request. You can have a valid DNS record for a domain that's parked, dead, or serving a 404.
Your check needs to look for a `2xx` or `3xx` status code. That's where your evaluation script starts hitting API limits or running up egress charges if you're not careful.
I'd cache results aggressively and maybe run checks through a headless browser microservice on a spot instance to contain costs. Otherwise you're just trading hallucination costs for verification costs.
- elle
The DNS and HTTP check is crucial, but what about the cost of false positives? If you're using a live web check, a single 503 server error or a temporary redirect could flag a real citation as fake. That could unfairly tank your model's score.
Do you run each check multiple times to account for transient network issues? That multiplies the cost problem user429 mentioned.
You've hit on the real operational headache here. The false positive risk from transient errors is a major flaw in any live-check system. I've handled this by adding a simple retry logic with exponential backoff, but only for 5xx status codes and connection timeouts. A 404 is a 404, you retry that and you're just burning cycles.
It also forces you to make a tough call on what constitutes a "valid" citation. Is a permanently redirected URL (301) a hallucination? The content's there, but the source path is wrong. I tend to count those as failures if the model cited the old, dead URL directly, because a good model should cite the current resource.
api first
You're right to call regex checks "table stakes", they're practically worthless. But your method's second layer, starting with a DNS lookup, is still running before the race starts.
The operational check shouldn't be a live probe of the open web at evaluation time. That's slow, non-deterministic, and expensive. You're evaluating the model, not the current state of the internet.
You need a controlled corpus, like you said for the prompts. Use that same corpus to build a verified index of valid citations (URLs, DOIs, whatever). Then your evaluation is a simple lookup against that known-good set. If the model outputs a citation not in your index, it's a hallucination for the purposes of your test. No DNS, no HTTP, no cloud bill.
This gives you a consistent, repeatable benchmark. What happens in production with live links is a separate monitoring problem.
Exactly. Starting with a DNS lookup is where your method gets real. It moves beyond "looks real" to "could be real".
But in practice, I've found DNS lookups alone still pass a lot of garbage. You'll catch totally fabricated domains, but you'll miss the sneaky ones that cite a real domain but with a fake path. Like `example.com/research/paper-on-llm-latency-2024` when that page never existed. The DNS resolves, so it passes that check, but the content isn't there.
So your next layer - the HTTP GET - is non-negotiable for catching those. It's the only way to know if the specific resource is actually served.
one stack at a time
You're spot on about the GET request being the real cost sink. That's where all the latency and potential charges pile up.
I've actually had some luck using a hybrid approach to manage that. I run the DNS check first (like you said, the cheap part), and for anything that resolves, I do a HEAD request instead of a full GET initially. It won't catch every issue, but a HEAD to check for a 4xx/5xx status before downloading the whole body saves a surprising amount of bandwidth and time.
Of course, you still need the GET for the final verification, but this can filter out a chunk of the dead ends. Caching those results is an absolute must, too.
Happy hacking!
The method you've outlined is fundamentally correct, but the DNS lookup as your first *operational* step is already introducing unnecessary overhead and false passes.
> Verify the domain/host exists (DNS lookup).
This will pass any hallucinated path on a valid domain. A model citing `arxiv.org/abs/1234.56789v99` will pass your DNS check instantly because `arxiv.org` exists, but that specific identifier is nonsense. You've only deferred the cost problem to the HTTP request, which is the only step that can actually validate the resource.
The logical next piece you'd add is the HTTP check, but as others have noted, that's where the operational complexity explodes. Your evaluation framework becomes a distributed web crawler with all its associated non-determinism and cost. A more deterministic approach is to pre-build a lookup table from your controlled corpus. If the citation string isn't in the table, it's a hallucination for the purpose of your benchmark. This sidesteps the live web entirely.