For the last six months, I've been tracking every suggestion Grammarly made on my technical blog posts. I wanted to see if it was truly helpful for our kind of content—full of code snippets, commands, and niche terminology—or if it was more of a distraction. I used a simple markdown-based workflow to capture the data.
My process was straightforward:
- I wrote all drafts in VS Code with the Grammarly extension.
- For each post, I exported the Grammarly "Performance" score and suggestions into a separate log file.
- I categorized each correction I accepted or dismissed.
Here's a snippet of the log structure I kept:
```yaml
Post: "Managing Pod Disruption Budgets in K8s"
Date: 2023-10-15
Initial_Score: 78
Final_Score: 92
Accepted_Suggestions:
- Punctuation: 12
- Conciseness: 3
- Incorrect spacing: 5
Dismissed_Suggestions:
- Tone adjustments (too formal): 7
- "Word choice" on technical terms (e.g., "HPA"): 4
- Incorrect comma rules for clauses: 2
```
**Key Findings:**
* **It excels at mechanical fixes:** Catching missing articles, repeated words, and straightforward punctuation errors. This is its strongest value-add, saving a final proofreading pass.
* **It struggles with technical context:** It frequently flagged correct Kubernetes resource names (`ConfigMap`, `DaemonSet`), common commands (`kubectl apply -f`), and even GitOps as "spelling errors" or suggested unnatural rephrasing.
* **The "clarity" and "engagement" suggestions often miss the mark:** Our writing needs to be precise, not necessarily "conversational." I dismissed nearly 80% of suggestions in these categories because they introduced ambiguity or softened definitive statements crucial for tutorials.
* **Score is a poor metric for quality:** A post full of correct technical terms lowers the initial score dramatically. The final score is only high if you accept all generic suggestions, which can dilute the technical voice.
**My Adjusted Workflow:**
I now use it as a **late-stage linter**, not a co-writer.
1. First draft is written without any grammar tool.
2. Technical review passes happen with peers.
3. **Only then** do I run Grammarly, with all "style" suggestions (`Engagement`, `Delivery`) disabled in settings.
4. I accept only the mechanical corrections (`Correctness` category) and carefully review any "clarity" suggestions.
For our community, it's a useful tool if you cage it properly. Out of the box, its defaults will fight you on terminology and style. The data shows it catches enough simple errors to be worthwhile, but you must be the final authority—it doesn't understand `terraform plan` or `istioctl analyze`.
The log structure is the interesting part. I've seen teams try to build entire linting pipelines around similar YAML outputs, only to realize they're just recreating spellcheck with extra steps. Did you track whether the "performance score" actually correlated with any meaningful metric, like reader engagement or reduced support questions? I've found those vendor scores are usually calibrated for generic business writing and get completely unhinged around technical terms.
Your point about it struggling with technical terms is why I turned it off entirely for docs. It kept insisting "Kafka" should be "café" and that "idempotent" was a misspelling. The mechanical fixes are useful, but at that point, I'd rather run a focused dictionary and a simple passive voice detector in the CI step. That way you're not fighting the tool's desire to make your Helm chart read like a company newsletter.