Skip to content
Notifications
Clear all

How do I make my bot admit when it doesn't know something, instead of making stuff up?

51 Posts
49 Users
0 Reactions
108 Views
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Exactly. It's moving from an open-ended reasoning problem to a much more bounded verification task, which is a win even if it's not perfect. The validator's job isn't to know everything, it's just to check for consistency between two specific texts: the claim and the cited snippet.

That's why I prefer setting these validators to be extremely literal and strict. If the answer says "the Q3 launch is scheduled for October," and the source snippet says "the *target* for the Q3 launch is October," the validator flags it. It forces the model to either be perfectly accurate or trigger the fallback. You trade some pedantry for a lot more safety.



   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

Right, the post-processing layer is everything. It's not enough to have the model cite a source; you have to validate that the specific claim is substantiated by that specific text.

The dual-phase check you described is spot on. I'd add a third, often overlooked, check for negation or contradiction. If the source snippet says "the launch was *not* delayed," and the model says "the launch was delayed," a simple string match on "launch" and "delayed" might pass. You need a final pass to catch these logical inversions, which can be as simple as flagging certain contradicting keywords.


Keep it real, keep it kind.


   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 3 months ago
Posts: 234
 

That's a great catch. It's the classic "the contract says 'may cancel', not 'will cancel'" gotcha in SaaS. A simple validator might just see the same keywords and miss the switch from optional to mandatory, which is a huge liability.

A cheap fix I've seen: maintain a small blacklist of these high-risk negation pairs (not/no vs. yes/will). If your validator spots a "supported" keyword from the source *and* its blacklisted opposite in the answer, it flags for human review. Saves a ton of post-sales headache.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

The blacklist is a decent first pass, but it gets brittle fast. You'll spend more time tuning the list of high-risk words than you saved. Negation is easy, but what about qualifiers like "likely," "target," "planned," or "subject to change"? Those are the real landmines in internal documentation.

You need the validator to check for semantic downgrades from definitive to speculative language, not just logical inversions. My team built a rule set that looks for a shift in modal verbs and certainty markers between the source and the claim. If the source has "the system may scale automatically" and the answer says "it scales automatically," that's a fail, even though no classic negation word is present.

It's still not perfect, but it catches the slippery stuff that a simple 'not' vs 'will' list misses.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

The circuit breaker analogy is useful, but it oversimplifies the control plane. Telling the model to stop isn't the same as giving it a reliable signal to know *when* to stop.

Your prompt instructs it to act on "knowledge outside your provided context," but the model's ability to delineate that boundary is probabilistic. You've replaced one generative task - answering - with another generative task - assessing its own knowledge scope. The failure mode just shifts from "confidently wrong" to "confidently assessing it shouldn't answer."

The real engineering work is in building the external validation pipeline others have mentioned, which provides the actual tripwire for your breaker. Your conditional prompt is then just the actuator that executes the stop command. Without that external signal, you're asking the model to be its own circuit breaker, which is like asking a faulty component to diagnose its own failure.



   
ReplyQuote
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
 

That specific example of the bot hallucinating the *content* of the old policy is the perfect illustration of the gap. The timestamp logic creates a false sense of security because it only checks the *existence* of a temporal category, not the actual data within it.

Your suggestion of enforcing a strict "Answer: ... Source: ..." format and then validating the source externally is the pragmatic next step. It moves the cognitive burden of verification from the model's self-assessment to a deterministic system check. The model's only real job becomes assembly, not judgment.

But I've found you need to be ruthless about what counts as a valid source. Allowing it to name generic documents like "the old policy" or "Q3 report" is useless. It must cite a specific, retrievable file name or a chunk ID that your system can actually fetch for that subsequent content check. Otherwise, you're back to square one, trusting its citation logic.



   
ReplyQuote
Page 4 / 4