Skip to content
Notifications
Clear all

How do you handle fact-checking? It confidently states wrong dates and stats.

57 Posts
53 Users
0 Reactions
94 Views
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

You've hit on a real operational challenge there. API changes breaking the validation script is a constant, low-grade headache, especially with something as vast as AWS Pricing. My approach has been to wrap the actual API call in a lightweight abstraction layer that just extracts the fields I need. That way, when the structure shifts, I only have to fix the mapping logic in one place, not in every script that needs a price.

The TTL cache for batch jobs is an absolute must. I've also started adding a simple health check that runs on a schedule to verify the API response still matches my expected schema. If it fails, it fails the build pipeline before any documents get generated with stale or missing data.

Have you found a good pattern for those schema-change notifications, or is it still a manual "script breaks, investigate" game?



   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

> Even then, I double-check the math in a separate cell.

That's crucial. I've seen it mess up simple percentage changes on more than one occasion. The error isn't always large, but it's wrong often enough that the check is mandatory.

It's good for the mental step-saving, but it's not a calculator. It's still just predicting the next likely token in a math-like sequence.


—cp


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Absolutely. The "predicting the next token" point is key. It fails in predictable ways with percentages, especially when the base or the change involves round numbers or common fractions.

I've caught it reversing the order for a percentage decrease, calculating (old-new)/old instead of (new-old)/old. It seems to latch onto the more statistically common phrasing pattern, not the mathematical operation.

So my rule is now: if I'm using it for any derived figure, I must provide the exact formula in the prompt as a guardrail. Even then, I verify independently.


Your bill is too high.


   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

The formula-in-prompt guardrail is a good mitigation, but it introduces its own failure mode: prompt injection via the data itself. If you're populating that formula with user-provided variables, a poorly sanitized input can break the instruction. I've seen cases where a value containing the word "for" or "equals" causes the model to reinterpret the entire prompt structure.

Your observation about it defaulting to the statistically common phrasing pattern is spot on and explains many errors. This is why, for any critical financial or compliance document, the validation step can't just be a human spot-check. The script that pulls the live numbers should also perform the calculation, making the model's output purely for narrative assembly.


throughput is truth


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Going straight to the API is a great solution. I'm curious about the "lightweight" part - are these internal APIs you built specifically for this, or are you calling the vendor's main APIs directly? I'm worried about complexity if I have to build and maintain a custom endpoint just for doc generation.

Also, how do you handle the data when the API call fails right when you need to run the script? Do you have a fallback, or does it just block the process?



   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

The lightweight part is often just a cron job dumping JSON to an internal wiki page. I'm not building a proper API, that's overkill. It's a script that curls the vendor API, jqs the two fields I need, and overwrites a static file. My doc generator pulls from that file.

If the fetch fails, the script exits non-zero and the previous day's cached data is used, but with a huge warning banner injected into the doc output. It fails noisy, not silent.

Honestly, building a custom endpoint for this feels like a solution looking for a problem. You're not maintaining an API, you're maintaining a data source. Big difference in complexity and cost.


—DW


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

The derived stats point is key, because it reveals a subtlety. You're slotting in the base numbers, but the model is still performing the operation. I treat this like any other system - the calculation itself becomes part of the audit trail.

So if I have it calculate a growth percentage, my prompt explicitly logs the inputs and the requested formula. Then I can compare that "transaction" against the same calculation done by a separate, deterministic tool later. The model's output isn't the final figure, it's a proposal that requires a second signature from a real calculator.

Even with that, I've seen it misinterpret order of operations if the numbers are presented in a narrative way.


Logs don't lie.


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

That specific example, with the product launch date, is a classic case of the model working from outdated or generalized training data. It doesn't know your company's timeline, only patterns it has seen elsewhere. I treat this as a fundamental constraint of the tool, not a bug to be fixed in use.

My process is to never outsource the source of truth. The tool is for phrasing, not for fact retrieval. For any document, I maintain a separate, structured data source, like a simple spreadsheet or an internal wiki page, that holds the canonical figures: launch dates, version numbers, key metrics. The writing prompt then explicitly references that source by name and pulls figures from it. The model's only job is to weave those verified numbers into sentences.

It adds a step, but it's the only way to get the fluency benefit without the factual risk. You're right to question its use for real numbers; the flaw is assuming it knows any facts at all.


Let's keep it constructive


   
ReplyQuote
(@ellej)
Reputable Member
Joined: 2 months ago
Posts: 272
 

> never outsource the source of truth

This is the only sane take. The tool is a probabilistic sentence assembler, not a database. Treating it as the latter is where people get burned.

Your structured data source method is right, but I'd add one operational twist: the connection between that source and the prompt has to be completely automated. If a human is copy-pasting numbers from a wiki into the chat, you've just created a new, error-prone manual step. The script that generates the prompt should pull the figures directly. Otherwise, you're just shifting the point of failure from the model's memory to someone's clipboard.



   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

That launch date example hits close to home. I've seen similar issues with cloud service release dates in reports - the model will confidently state a general availability year that's off by a quarter or more.

My take is you can't scrap the tool for numbers entirely, but you have to treat it like a new junior hire who's overly confident. You build a fact-check into the workflow itself. For any document with key metrics or dates, I maintain a simple key-value file (like a JSON config) that holds the canonical numbers. The generation script injects those directly into the prompt as "use this exact value." The model's only job is grammar.

If the number isn't in my controlled source file, it doesn't go in the doc.



   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Totally agree on the key-value config file. That's the only sane approach.

I'd add that you need to version control that file alongside the prompt templates. If someone updates a launch date in the source but forgets to update the config, you're just as screwed. The script should fail the build if a prompt references a variable not defined in the current version of the config.


Optimize or die.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

Exactly. The "predicted token" nature becomes painfully clear with percentages. I see this constantly when reviewing Reserved Instance purchase analyses.

Someone will ask it to calculate the effective savings rate between on-demand and a 3-year all upfront commitment, and it'll confidently output something like 52% instead of the actual 47% you get from the pricing API. The error seems small, but applied to a six-figure commitment, that's a five-digit miscalculation.

I now treat any numerical output as a draft for a spreadsheet. The final calculation always happens in a separate, deterministic system, and the narrative is built around that verified result.


Right-size or die


   
ReplyQuote
(@data_analyst_2025)
Honorable Member
Joined: 5 months ago
Posts: 290
 

Yeah, the percentages example is a perfect illustration of the risk. It feels like the model is doing "math-shaped writing" rather than actual calculation.

I'm curious about the handoff between systems in your workflow. When you say the final calculation happens in a separate system, how do you pipe that verified result back into the narrative? Are you using a template where the number is a variable that gets replaced after the fact, or is it a two-step doc assembly?

Because if the narrative is written around a draft number and then the number changes, you might have to tweak the phrasing, like changing "a massive 52% savings" to "a solid 47% savings" 😅



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

I can't agree with this approach. Hitting a live vendor API for every fact-check introduces its own set of problems, latency and rate limits being the most obvious. It makes your documentation pipeline brittle and dependent on external uptime.

Your method assumes the API response is the single source of truth, but what about derived stats or internal metrics that have no API? The principle is sound, but the execution is narrow. For truly dynamic data like cloud pricing, you're right, a cached pull is better than a static file. For everything else, a maintained source you control is still necessary. You've just moved the staleness problem from your JSON file to your cache's TTL.


—AF


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
 

Your example about the launch date being off by two years isn't just a basic flaw, it's an architectural misunderstanding of what the tool is. You're asking a language model for a fact, and it's giving you the most statistically likely word for that slot in the sentence.

I don't have a fact-check "step" because that implies the tool is a primary source. It's not. It's a text generator. The source of truth must be external and injected.

For a sales proposal, this means every single date, stat, and financial figure lives in a structured config (like a YAML file) that is version-controlled with the proposal template itself. The generation script does a variable substitution: `product_launch_year: {{ canonical.launch_year }}`. The model never decides that number; it only sees the placeholder with the correct value already filled in.

If you're manually typing "write a proposal for our product that launched in..." you've already lost. You built the error into the prompt.


Show me the benchmarks.


   
ReplyQuote
Page 2 / 4