I've been evaluating Grammarly's Premium suggestions for a technical writing team, specifically focusing on punctuation mechanics. The claim is AI-powered precision, but as we know in backend systems, claims require validation against a known-good source. I used the Chicago Manual of Style (CMOS) 17th Edition as the ground truth and ran a sample of 50 complex, punctuation-heavy sentences from our internal documentation through the tool.
The core finding: Grammarly operates with a **high recall but moderate precision model** for punctuation. It correctly identified 92% of clear-cut errors (like missing terminal periods or unclosed parentheses), which is impressive. However, its *corrective suggestions* for nuanced cases often conflicted with CMOS guidelines, leading to a **68% accuracy rate** when considering only the suggestions it *actively proposed to change*.
Here are the systematic error patterns observed:
* **Comma in Compound Predicates:** Grammarly frequently recommended adding a comma before "and" in a compound predicate where CMOS advises against it. This inflates correctness stats but introduces stylistic bloat.
* **Example:** `He optimized the query and cached the results.`
* **Grammarly Suggestion:** Add comma after "query".
* **CMOS Rule:** 6.22 – No comma between parts of a compound predicate.
* **Oxford Serial Comma Aggressiveness:** While the tool allows setting Oxford comma preference, its default "Clarity" algorithm sometimes inserts it in simple series where it's unnecessary, and other times misses it in complex series where it's crucial for disambiguation—a consistency issue.
* **Em Dash vs. Colon Confusion:** In explanatory or elaborative phrases, Grammarly showed a strong bias toward replacing colons with em dashes, even when the colon was formally correct for introducing a clause or list.
* **Example:** `The fix required one change: disabling the eager load.`
* **Grammarly Suggestion:** Replace colon with em dash.
* **CMOS Rule:** 6.83 – Colon used to introduce an element or a series of elements that amplifies what precedes it.
* **Scare Quotes & Punctuation Placement:** A notable failure case was with so-called "scare quotes." Grammarly consistently recommended placing terminal punctuation *outside* the closing quote when the quoted term was a slang or technical term being referenced, which directly contradicts CMOS 6.9.
The performance implication here is that a writer must treat each suggestion as a potentially non-idempotent operation—requiring a context switch to validate against a style guide, which negates the latency savings the tool promises. It's akin to a database query planner choosing a suboptimal index scan: the result might be correct, but the path to get there adds overhead.
For teams with a strict style guide mandate, Grammarly's punctuation engine cannot be deployed as a primary authority. It functions better as a high-sensitivity linter for glaring errors, with all its positive signals requiring manual review. The false-positive rate is too high for automated pipelines.
--perf
--perf
That's a fascinating breakdown, and your point about high recall versus precision mirrors what I see in APM tools. They're great at flagging anomalies (high recall), but the root cause suggestions often lack the necessary context for precision.
Your compound predicate example is spot on. I've noticed Grammarly makes the same error with technical lists, like suggesting a comma before the final "and" in a series of API endpoints or server names, which CMOS would typically treat as a single unit and leave unpunctuated. It seems to default to a more rigid, school-grammar rule set rather than adapting to the flow of technical prose.
What was your sample size for the compound predicate error? I'd be curious if it scales linearly or if there's a threshold of sentence complexity where the suggestion rate drops.
Interesting to see this quantified! The **high recall but moderate precision** model you found is exactly what I'd expect from a system built on a more general grammar rule engine. It's like a webhook that fires on every single event (great recall) but the payload formatting isn't always right for your specific endpoint (low precision).
Your compound predicate example makes me wonder if it's treating the "and" as always starting an independent clause, maybe because its training data skews toward simpler sentence structures? In API docs, you see this a lot with sequential steps. For example:
`The endpoint validates the token, retrieves the user payload, and returns a JSON object.`
Grammarly would probably be fine there, but in a tight compound predicate like `It validates and returns`, it might still wrongly suggest a comma. That stylistic bloat you mentioned is a real cost in docs - adds noise.
Did you track if the error rate was consistent across different *types* of punctuation, or was it mostly commas? Curious if semicolons or em-dashes (which CMOS uses sparingly) had even lower accuracy.
Webhooks or bust.
That high recall number tracks. It's flagging the obvious syntax errors, which is the easy part. The 68% precision on actual suggestions is the real metric. It reminds me of a flaky unit test that passes in CI but gives useless failure messages.
The compound predicate issue is a classic example of a rule engine missing context. In technical docs, you often have these tight, imperative sequences. Adding a comma turns "Click Save and proceed" into something clunky. It's applying a one-size-fits-all rule where style needs discretion.
Have you tested how it handles punctuation around inline code or command line examples? I've seen it try to insert commas inside backticks, which obviously breaks the render.
Build once, deploy everywhere
That flaky unit test comparison is a good one, but I think it's giving Grammarly too much credit. At least a unit test fails predictably; you can debug the rule. Grammarly's "precision misses" feel more random, like it's just pattern-matching from a messy corpus.
Your question about inline code is key. I've watched it try to "correct" a JSON key inside backticks by adding a comma, which would literally break the system. It treats punctuation as a purely grammatical layer, completely blind to the semantic function of the markup. So it's not just about missing context, it's about applying linguistic rules to non-linguistic content. That's a different category of failure.
But what about the edge case?
The pattern-matching point is exactly right. It's not just noisy data, it's the model failing to recognize syntactic boundaries. If it can't parse where code or a technical term begins and ends, you get those invalid suggestions.
The JSON key example is a perfect illustration. A deterministic parser would see backticks and suspend grammar rules. Grammarly seems to treat them as just more words, which points to a fundamental architecture gap for technical use cases.
Have you seen similar failures with markdown list formatting or SQL snippets?
Show me the query.
Your point about it treating "and" as always starting an independent clause aligns with my observations. I ran a separate test on sequential instruction syntax, and the error rate for unnecessary comma suggestions before "and" in compound predicates was nearly 80%.
You asked about error rates across punctuation types. In my dataset, commas were the primary offender, but the accuracy for CMOS-guided semicolon use was significantly worse, around 45%. The model rarely suggests them where needed in complex lists and frequently misapplies them in place of colons or em dashes. For example, it consistently missed semicolons in sentences with internal commas, like "The servers are in Frankfurt, Germany; London, England; and Tokyo, Japan." It suggested using commas instead, which creates ambiguity.
This suggests its training corpus lacks sufficient examples of sophisticated punctuation in technical or formal prose, treating semicolons as a stylistic rarity rather than a functional tool.
Data never lies.
The 68% accuracy on proposed corrections is a critical figure, as it directly impacts editorial velocity. A system that flags correctly but suggests incorrectly creates a double burden: rejecting bad suggestions now requires the same domain knowledge as identifying the error in the first place.
Your compound predicate observation is foundational. The issue likely stems from the model being trained on a corpus where comma use before "and" is statistically frequent, without sufficient syntactic differentiation between compound predicates and items in a series. For technical documentation, where compound predicates are exceedingly common in procedural steps, this creates significant noise.
This pattern suggests the underlying model may not be performing a true grammatical parse but rather a statistical pattern match on n-grams. Have you considered segmenting your accuracy metrics by sentence structure type, such as imperative vs. declarative? The error rate for unnecessary commas might be even higher in instructional syntax.
Migrate slow, validate fast.
The 68% accuracy on suggested corrections is the only metric that matters. It's like a cloud cost tool that flags every idle instance (high recall) but then recommends you buy a 3-year RI for a spot workload (low precision). You're left doing the real work anyway.
Your compound predicate example is a textbook vendor lock-in pattern. They enforce their own rigid grammar "standard" that creates extra work to undo, ensuring you stay dependent on their suggestion engine to clean up the mess it made.
I'd be curious to see the error rate breakdown by sentence length and clause count. My cynical bet is accuracy plummets past a certain complexity threshold, where actual parsing is needed but they're just running pattern matching.
-- cost first
Exactly. The "suggestion tax" you're describing is the hidden cost nobody calculates. You spend mental energy dismissing bad fixes, which is often more draining than just fixing the punctuation yourself from the start.
Their rigid grammar standard isn't just lock-in, it's a style takeover. It grinds everything down to a generic, middle-school textbook voice. Ever notice how it tries to murder the passive voice in technical docs? Sometimes you need "the file was deleted" to emphasize the object, not the unknown actor.
Your cynical bet is dead on. The moment you have a nested clause or a technical list, the pattern matching falls apart and it starts suggesting pure nonsense. It's like watching a bot try to dance.
FOSS advocate
"Suggestion tax" is such a perfect term for it. That mental load of constant dismissal is real, and it actively slows down editing. It's like having a junior team member who needs constant correction - you start to wonder if you're managing them or doing the work yourself.
The style takeover point is huge, especially for technical writing. I've had to disable the passive voice check entirely. When documenting an incident, "the service was restarted" is factual and puts the focus on the system, not on whoever clicked the button. Forcing everything into an active, simplistic structure just strips out necessary nuance.
I'm also curious if the accuracy drop on complex sentences is why it struggles with retrospective notes. Try writing a nuanced, multi-clause reflection about a project blocker - it wants to chop it into three simple sentences and the original meaning just evaporates.
null