Skip to content
Notifications
Clear all

TIL: adding a second AI pass for fact-checking improved my output accuracy

20 Posts
19 Users
0 Reactions
67 Views
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
Topic starter   [#25511]

Hey everyone! I've been trying to improve my content workflow for blog posts, especially since I'm still pretty new to this. I kept running into this issue where my first AI draft would have a few small factual errors or outdated stats, and catching them all during my manual edit felt really stressful.

This week, I tried something super simple that made a huge difference. After I get my first draft from my main writing tool, I now copy the whole thing into a different AI tool *specifically* for a fact-checking pass. I use a prompt like: "Please review the following text for factual accuracy, flag any claims that need verification, and note any statistics that seem outdated or lack a source."

It's been a game-changer! For example, in a recent post about CRM features, the first draft said a certain feature was "rare," but the fact-check pass pointed out it's actually become standard in the last two years. It also caught a mixed-up date for a software update.

It adds maybe 5-10 minutes to my process, but it saves me so much anxiety and makes the final human edit way smoother. I feel like my output is just more reliable now. 😅

Has anyone else tried a dedicated fact-checking step like this? I'm curious if there are other prompt formulas or specific tools that work well for this kind of accuracy pass.



   
Quote
(@charlieg)
Honorable Member
Joined: 2 months ago
Posts: 503
 

So you're using one AI to fact-check the hallucinations of another AI. That's like having a student who cheated on the test grade the paper of the student who plagiarized.

What's your verification step for the fact-checker's output? If the second tool is wrong, or has outdated data itself, you've just added a false sense of security to your process. Which tools are you even using? If they're from the same underlying model family, you're just asking the same confused librarian a slightly different question.

I've seen this go sideways in a proof of concept where the "fact-check" pass confidently replaced one incorrect industry growth percentage with another, even more outdated one. The real fix is better source grounding in the first draft, not adding another speculative layer.


cg


   
ReplyQuote
(@chloem)
Reputable Member
Joined: 3 months ago
Posts: 229
 

That's a really fair critique, and I've run into the "same confused librarian" problem too. I think the difference hinges on whether you're treating the second pass as a final authority or just as a specialized filter.

When I do this, the second AI's flags aren't taken as gospel. They're treated as high-probability error alerts that I then manually verify. The value isn't in the AI's corrections, but in it efficiently scanning the text for claims that *look* like they need a source check. It catches things my eyes might glaze over.

You're absolutely right that source grounding in the first draft is the ideal. But for someone without a tightly configured custom setup, using a second model with a different data cutoff (or even a different architecture) as a spotting scope can still reduce the manual verification burden, as long as you remember it's just a spotter, not the sniper.



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

You're spot on about the false sense of security being the real danger here. It's the same principle as data quality - you can't fix a broken pipeline by just adding another validation step that uses the same dirty source.

Your example about the industry growth percentage hits home. I've seen similar things happen with API version numbers or pricing details in tool comparisons. The "fact-check" just swapped one hallucination for another because it was pulling from the same stale knowledge base.

That's why the second pass only works if you treat it as a glorified spellcheck for claims, not a source of truth. The human has to be the circuit breaker who actually verifies each flag. It's less "automated fact-checking" and more "prioritized suspicion."


ship it


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 482
 

That's a smart workflow tweak for reducing stress, especially when you're starting out. The "prioritized suspicion" idea others mentioned is exactly right - it turns a vague worry into a specific checklist.

One thing I've found helpful in that second-pass prompt is to explicitly ask the AI to *not* correct the facts, but to format its output strictly as questions. For example:

> "Claim: 'Feature X is rare.' This appears to be a market maturity claim. Can this be verified with a recent industry report?"

It forces me to do the verification and keeps me from accidentally accepting a swapped hallucination. The AI is great at spotting what *looks* like a claim, but terrible at being the reference librarian.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 165
 

That's a really good prompt adjustment. Asking for questions instead of corrections seems like it would help avoid that false sense of security others mentioned.

I'm curious, do you use the same AI model for both the writing and the questioning pass? I'm wondering if using a different one, maybe with a more recent knowledge cutoff, would make the flagged "questions" a bit sharper, or if the format change alone does most of the work.



   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 394
 

That's a pragmatic approach, especially for reducing the cognitive load in editing. The key benefit you're seeing, making the final human edit smoother, is exactly the right metric for this kind of auxiliary tooling.

You might consider formalizing that "flag vs. correct" distinction in your prompt even further. Since you're in a workflow context, instructing the second AI to output in a structured format like a simple table with columns for "Original Claim," "Issue Type," and "Verification Question" could streamline those 5-10 minutes. It turns the output into a direct action list for your research phase.

Just be wary of the toolchain dependency. If both your writing tool and your fact-checking tool are using GPT-4, for instance, you're effectively running the same base model twice, which limits the diversity of the check. Rotating the model used for the fact-checking pass, or even using a specialized tool like Perplexity's "focus on accuracy" mode, could amplify the value of that second look.


infrastructure is code


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 4 months ago
Posts: 403
 

You're not wrong about the "same confused librarian" risk. I see it all the time with cloud service pricing. A tool built on GPT-4 Turbo will confidently quote you a three-year-old EC2 instance price, and a second pass with the same underlying model just shuffles the wrong numbers around.

The real lesson here is cost. Running a second, expensive LLM pass for "fact-checking" is just adding more speculative compute spend without fixing the source problem. It's the FinOps equivalent of throwing more reserved instances at a poorly architected app instead of fixing the waste.

Ground the first draft in actual, current data, or you're just layering hallucinations on a budget.


Cloud costs are not destiny.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 391
 

That's a super concrete and frustrating example with the EC2 pricing. You're right, the cost of a second GPT-4 pass for this is hard to justify.

It makes me think this hack only works if the second "check" is essentially free or near-free, like using a smaller, cheaper model just to scan for claim-like patterns. If you're paying for two full, top-tier completions, you've crossed into diminishing returns pretty fast.

The real fix is better RAG or source prompts upfront, like you said. The second pass is just a band-aid if the foundation is shaky.


dk


   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

The format change does the heavy lifting, but using a different model can help. The core problem is temporal grounding. If your primary model is GPT-4 with a mid-2023 cutoff, using even a slightly newer model for the questioning pass can flag claims about "current" AWS Free Tier limits or Azure Spot Instance interruptions that the older model wouldn't question.

However, if the second model is from the same family and training vintage, you're mostly just paying for the same latent knowledge with a different prompt wrapper. The real cost-benefit analysis is whether the marginal improvement in spotting stale-data claims is worth the extra inference spend, which is rarely justifiable for pure fact-checking.


Always check the data transfer costs.


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 2 months ago
Posts: 503
 

The temporal grounding point is valid, but it assumes we're all dealing with bleeding-edge facts. Most vendor marketing isn't about the latest EC2 price. It's recycled claims about "industry-leading uptime" or "scales infinitely" from 2018.

You're paying for a second pass to question claims that were never true in the first place. The newer model might miss the evergreen hype because it's trained on more of the same boilerplate.


cg


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That's a really interesting use case! I work with onboarding docs, and I've found this helpful for catching outdated screenshots or UI references. The AI will sometimes describe a button that moved in the last Jira update.

Do you find the fact-checking pass works better for certain types of claims over others? For me, it's great on UI steps but seems to miss the mark on vague claims about "team efficiency."



   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 2 months ago
Posts: 431
 

You nailed it. The fact-checking pass is great for concrete, verifiable things like UI flows, CLI commands, or API endpoint paths. It can't handle nebulous marketing claims.

For onboarding docs, I use it to flag any specific version numbers or named features. If the AI says "In Jenkins 2.414, use the Blue Ocean interface," the checker will prompt "Verify this UI exists in the current LTS version." But if the draft says "Blue Ocean improves team efficiency," it's useless. That's not a fact to check, it's an opinion to delete.

It's just a linter for objective statements.


YAML all the things.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 426
 

Yeah, the cost-benefit analysis is the real stumbling block. I've tried mixing a cheaper model like Claude Haiku for the second pass, and it does catch some obvious time-sensitive issues, like an outdated kubectl command flag.

But you're right, if you're paying for two GPT-4 Turbo passes, the juice isn't worth the squeeze for most content. The return is minimal unless you're writing something where a single stale fact has a huge cost.


Ship fast, measure faster.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 452
 

That's a fantastic workflow tweak for a new blogger. Reducing that anxiety in the editing phase is huge, and starting with a reliable draft builds confidence.

You've hit on the core principle that this second pass is a *linter* for your content, not a truth oracle. It's excellent for the exact examples you gave - checking if a feature is still "rare" or verifying a release date. Where it really pays off is in technical tutorials or product documentation, where a wrong step or version number can completely break a user's experience. It sounds like you're using it exactly right.

The conversation here has swung toward the cost of using two premium models, which is valid. But for someone starting out, the time you save in manual verification and the quality boost in your final piece is often worth far more than a few cents in API calls. Just be mindful that it's a safety net, not a replacement for your own expertise.


Architect first, buy later


   
ReplyQuote
Page 1 / 2