Alright, let's get this one on the record. I've been running Jasper through its paces for a few months now, primarily for generating marketing copy and blog outlines. There's a persistent and frankly dangerous habit I can't seem to break: it keeps inventing statistics.
I'll ask for "a paragraph on the benefits of email marketing for small businesses." It'll spit back something like "Studies show that 73% of small businesses see a 40% increase in customer engagement within the first quarter of implementing a strategic email campaign." That's a very specific, very confident pair of numbers. They are, of course, completely fabricated. There is no cited study, and those figures smell like pure simulation.
I've tried being explicit in my prompts: "Do not include any statistics unless they are commonly cited and verifiable." The result? It either ignores the instruction, or it gets cagey and uses vague weasel words like "many" or "a significant percentage," which isn't much better. It's swapped one form of nonsense for another.
This isn't just an annoyance; it's a data quality issue that makes the tool unusable for anything requiring factual rigor. If I wanted to A/B test copy, I can't have one variant leaning on phantom numbers. My attribution models are complicated enough without AI-generated fairy tales.
So, my question to the community: has anyone found a reliable prompt engineering technique or a workflow that forces Jasper to either cite real, known sources or, better yet, just not make up numbers at all? Or is this a fundamental flaw in how its model completes patterns, making it inherently untrustworthy for any statistically-adjacent content?
Data skeptic, not a data cynic.
Yeah, this is the core tension with these "brand voice" oriented tools like Jasper, isn't it? They're essentially fine-tuned on marketing copy - a genre where plausible-sounding, unsourced stats are weirdly the norm. The model learns that pattern and can't help but generate it, because for its training data, that *was* the "correct" output style.
You mentioned it swaps to vague weasel words when you push back. That's actually the model hitting its limits - it knows it shouldn't make up a number, but it has no real database of verified stats to pull from, so it retreats to the safest, most hollow phrasing. You're trading a concrete lie for a vague truth, which is often worse for copy.
A brutal but effective prompt tweak I've used is to explicitly forbid the genre's cliches: "Write a paragraph on [topic]. Do not use any statistics, percentages, or phrases like 'studies show,' 'research indicates,' or 'it is estimated that.' Write only about concrete, generally accepted benefits." It forces it into a different narrative mode.