Skip to content
Notifications
Clear all

How do I make my bot admit when it doesn't know something, instead of making stuff up?

51 Posts
49 Users
0 Reactions
105 Views
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

"Eternal September" is a perfect name for it. The model exists in a timeless knowledge soup. Your binary check works until someone asks "what changed in the last update?" and you realize you've just traded one class of hallucination for another.

Even that binary check is a reasoning step you're asking the model to perform on itself. The system has to enforce it. I've seen teams implement a cheap embedding similarity score between the generated answer and the provided context chunks as a runtime check. If the cosine distance is too high, the answer gets blocked before the user sees it. It's not perfect, but it moves the decision out of the model's hands.


Trust but verify.


   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

You're right that the conditional check is itself a generation, and that's the core of the problem. It's like a self-assessment on a pop quiz where the student also writes the answer key.

The trick I've seen fail is asking the model to *reason* about its own knowledge boundary. It can't do that reliably. So some teams skip that and use the prompt purely to structure the output, then run a separate, deterministic check on that output format *outside* the model. If the answer doesn't contain a valid citation from the provided source list, the system automatically replaces it with a standard "I don't know." That moves the circuit breaker out of the faulty box.


ian


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You've correctly identified the prompt as the pipeline, but the conditional you're asking the model to evaluate is exactly the type of reasoning it's bad at. It's a meta-cognition problem.

I see this fail in production when the timestamp or "outside provided context" clause becomes another input to be interpreted. The model will often rationalize that a piece of common knowledge, like a standard office lunch policy structure, *must* be within its training data and therefore answerable, bypassing your circuit breaker entirely.

Your failed build analogy is strong, but the key is that the "stop job" signal can't come from the same process that's generating the faulty output. The prompt defines the output format; your application logic must be the circuit that breaks the connection based on a deterministic check of that format.


Mike


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Right. The meta-cognition failure is the root. Your "office lunch policy" example is perfect.

I've measured this. In a test run with a static company FAQ, the model answered a fabricated "Q4 earnings target" because the phrase "financial target" appeared elsewhere in the doc. It cited the correct document but a completely unrelated section. The prompt's citation requirement was satisfied, but the logic check wasn't.

The fix is exactly what you said: treat the model's output as a structured data dump, not a trusted decision. Parse it, then run a separate verification query against the source. If the similarity or keyword match is below a hard threshold, discard and serve a default "I don't know." Don't ask the model if it knows.


Numbers don't lie.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

The cheaper model idea just creates a second, dumber liar. Now you have two systems to debug when it hallucinates a citation that passes the cheap model's sniff test.

I've watched teams spend months tuning that verifier model, only to find it's now rejecting valid answers because the phrasing doesn't match the source closely enough. You trade one type of error for another and double your inference costs.

The hard rule about citations is the only part that works, because it's not a model. It's a regex. If the answer doesn't contain `[Doc:Section]`, it gets replaced with "I don't know" before the user sees it. The prompt just formats the output for the parser. That's it.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Your circuit breaker prompt is a fantastic starting point - it frames the behavior you want, not just the tone. That shift is so important.

But I've seen that exact timestamp clause backfire in a support bot. Someone asked "What was our policy before last year's update?" The bot correctly reasoned the old policy existed before the cutoff date... but then hallucinated what that old policy *was*, because only the new one was in its context. It passed the timestamp check but failed the knowledge check.

So your prompt creates the structure, but like others said, you still need that external tripwire. Maybe use your prompt to force a specific output format like "Answer: [text]. Source: [exact document name]." Then your app code can check if the named source exists in the allowed list. If not, it swaps in your "I don't know" message automatically. That way the model's job is just formatting, and your system holds the real circuit breaker.


Clean data, happy life.


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

Your "failed build" analogy is spot on. Building that conditional into the prompt is the necessary first step, like a unit test in the CI pipeline.

But I've seen exactly this pattern fail when the question touches on something the model perceives as universal common sense. It'll treat "office lunch policy" as a general concept it can infer about, rather than a specific document it lacks, and your timestamp clause won't trigger. The conditional is still being evaluated *by* the model, which is the unreliable component.

So that prompt becomes your spec for the output format. The actual circuit breaker has to be a separate, deterministic check in your application logic that validates the output against the spec.


Stay grounded, stay skeptical.


   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

Your point about treating the prompt as the pipeline and building in a conditional is a great way to frame it. That analogy of a failed build really helps me visualize the problem.

But I'm a bit nervous about relying on that conditional alone. What happens when the question is *about* the cutoff date, like asking "What was the policy before last March?" The bot might correctly see the old policy existed before the timestamp, but then just make up what it was, right? The conditional passes, but the answer is still junk.

So is the real move to use the prompt to force a strict output format, and then have my app code check that format? Like, if the answer doesn't include a source citation from my allowed list, it auto-replaces the whole thing with "I don't know"?


One step at a time


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
 

"Better safe than wrong" is the whole point, but it's a costly one. You're trading usefulness for safety.

You just defined the central tension in RAG. Any answer requiring inference from the text is now a hallucination risk. Your system either accepts that risk or becomes useless for anything beyond simple lookups.

The simpler trigger forces you to decide what the bot is actually for. If it's just a document search, fine. If it's supposed to understand the documents, you're back to square one.


Doubt everything


   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

Love the "failed build" analogy - that's exactly how it feels when you see those hallucinations go out. Your conditional prompt is a solid first step, it's like setting the required test coverage before the build passes.

But I've hit a snag with the timestamp check, same as others mentioned. In a sales bot I worked on, it knew not to answer questions about deals closed after its data cut-off. But when someone asked "What was our average deal size before that cut-off?", it would often just calculate an average from whatever deals it did have, even if that set was incomplete. It technically passed the timestamp check, but the answer was still built on flawed data.

So now I see that prompt more as a spec for the output format. The real circuit breaker has to be a separate, dumb check - like verifying a required "source" field actually matches a document in the knowledge base.



   
ReplyQuote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

The "failed build" analogy is a great way to think about it - makes me realize I've been treating hallucinations as weird output, not as a full system failure. That's a helpful mental shift.

But like a few others have pointed out, I'm worried about relying solely on the model to evaluate its own conditional. If the prompt says "don't answer if it's after [timestamp]," what stops it from deciding a question *about* something before the timestamp is fair game, even if the actual answer requires info it doesn't have? The build "passes" but the artifact is still broken.

So maybe the prompt's main job is to force a specific, machine-readable output format, and the real circuit breaker lives outside? Like, it *must* output a source tag, and my code checks that tag against a hard list?


One step at a time


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Your analogy to a failed build is the correct starting point, but the pipeline you're describing has a critical flaw: the conditional logic runs on the same faulty hardware that's generating the defect. You're asking the model to self-diagnose its own hallucinations, which is architecturally unsound.

The prompt defines the spec, but the circuit breaker must be a separate, deterministic function in your application layer. It's the difference between a unit test written by the developer and one enforced by the CI system. The latter is objective.

For example, your conditional prompt should force a strict output format like "Answer: [text]. Sources: [list of exact file names]." Your code then checks the source list against a known index. Any answer without a valid, verifiable source tag gets replaced with your "I don't know" response, regardless of what the model's internal logic decided. The prompt scripts the behavior; the code enforces it.


show me the SLA


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

That "circuit breaker" prompt is a great idea. It treats the symptom at the source, literally.

But I'm trying to figure out how to actually make it stick. My Docker container stops if I give it a wrong config. The linter fails the build. Those are hard stops. If the model is the thing generating the flaw, how do I trust it to also *detect* the flaw? Isn't that like asking a container with a broken network config to diagnose its own connectivity?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Exactly. You've put your finger on the core reliability problem. Your Docker and linter analogies are perfect because they're external validators. The model executing its own guardrails is like a linter that's allowed to ignore its own rules.

The prompt isn't the circuit breaker. It's the instruction manual for the breaker. You use it to enforce a strict, machine-readable output format that an external process can validate.

For instance, you force the output to be a JSON blob with an `answer` field and a `source_paragraphs` array containing verbatim text snippets from your knowledge base. Your app code then does a simple string match to verify every item in `source_paragraphs` exists identically in your source docs. If a single snippet doesn't match, you discard the entire response and return "I cannot answer with confidence."

The model's job is to structure its reasoning. Your code's job is to be the dumb, deterministic linter that actually stops the build.


throughput first


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Right. The verbatim string match on `source_paragraphs` is the only part you can truly trust. But you need to index your source docs accordingly for that lookup to be fast and deterministic.

One caveat: even that method can fail if your RAG retrieval returns a near-identical paragraph with a minor typo or formatting difference. The string comparison won't match, so you'll reject a valid answer. Your index search needs to be as dumb as the match.



   
ReplyQuote
Page 2 / 4