It’s become a bit of a sport around here, hasn’t it? We’re all obsessively tracking perplexity, running human evals on “helpfulness,” and debating the perfect RAGAS score. Meanwhile, the actual output—the whitepapers, the battle cards, the competitive intel summaries churned out by our fancy AI sales enablement setups—is slowly becoming unreadable. It’s not *wrong*, per se. It’s just drowning in a thick, viscous soup of consultant-speak and empty calorific buzzwords.
I got tired of just complaining about it, so I built a simple tool to quantify the problem. I’m calling it the Jargon Density Index (JDI) tracker for now. The premise is laughably straightforward: you feed it a text block, and it measures the percentage of tokens that belong to a curated “offender” list. It’s not about technical terms, but about the filler language that signals a model has been trained on too many Gartner decks and corporate blog posts.
My initial lexicon includes, but is not limited to:
- **The Empty Verbs:** leverage, utilize, synergize, ideate, operationalize
- **The Vague Nouns:** landscape, ecosystem, paradigm, journey, space (as in “the AI space”)
- **The Redundant Adjectives:** holistic, robust, scalable, best-of-breed, mission-critical
- **The Prepositional Plagues:** “in order to,” “as it relates to,” “within the realm of”
Running a few dozen AI-generated “state of sales tech” whitepapers through it yields some depressing, if predictable, numbers. We’re routinely seeing JDIs between 8-12%. That means roughly one in every ten words is pure semantic fluff. The human-written ones from reputable sources? Clustering around 3-5%.
The immediate, practical application I see is in tuning prompts and evaluating models for sales enablement content. If your “personalized” battle card generator is spitting out something with a JDI north of 10%, you haven’t saved your sales team any time—you’ve just given them a slightly faster way to sound like a robot trying to sell to another robot. It’s a brutal litmus test for whether your AI output has any human spark left, or if it’s just elegantly reformatting the industry’s collective gobbledygook.
I’m now using it as a gate in our automated content pipeline. Any asset that trips a certain threshold gets flagged for human review and a prompt rewrite. The real question for this forum is: are we evaluating the right things? We chase metrics that look good in academic papers, but are we measuring if the output is actually *usable* by a stressed-out Account Executive trying to sound like a genuine person on a customer call?
I can share the basic script and the current offender list if anyone’s interested. It’s a blunt instrument, but sometimes you need a blunt instrument to break through the noise.
🤷