Skip to content
Notifications
Clear all

Sharing: My team's rubric for scoring AI-generated content quality

23 Posts
22 Users
0 Reactions
3 Views
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
Topic starter   [#29322]

Been forcing my team to use Notion AI for drafts. It's... fine. But "fine" doesn't cut it. We needed a way to score the output objectively, otherwise it's just vibes.

We built a simple 5-point rubric. Each point scored 1-5. It's brutal, but it works.
* **Accuracy & Hallucinations:** Does it make stuff up? 1 point if it invents stats.
* **Relevance & Focus:** Does it stick to the brief or go on a tangent?
* **Tone & Brand Voice:** Does it sound like us, or like a generic robot?
* **Actionability:** Does it have a clear next step or is it just fluff?
* **Originality of Insight:** Does it just rephrase the input, or add actual value?

Aim for 20+ total. Below 15, you're better off starting from scratch. Saves us from endlessly "polishing" AI slop. Try it.


CRM is a means, not an end.


   
Quote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

This is exactly what I was hoping to find. I've been trying to get our procurement team to adopt AI for drafting RFPs, and the feedback loop is always so subjective. Having a concrete score to point to would cut down on those "I just don't like the feel of it" debates.

One thing we've struggled with is the weighting. Is a critical hallucination on a fact just a 1 in Accuracy, or should it be an automatic fail for the whole draft? Your "below 15, start over" rule suggests the total score handles that, which is smart.

How do you handle disagreement on a score between team members? For example, if one person thinks the tone is a 2 and another thinks it's a 4, do you average it or debate until consensus? I could see that taking more time than the editing would.



   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

Love the rubric idea, especially the "below 15, start over" rule. That's a concrete decision point that saves so much time. I'm curious, do you find teams spend less time debating the score because they know a low total means a hard restart anyway? It feels like it would force more honest scoring.

How do you actually run the scoring? Is it a quick team huddle with the rubric on screen, or do people score independently first? Trying to picture the workflow.



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

I like the idea of a rubric, but I'm curious about the practicality. Scoring something like "Originality of Insight" feels super subjective to me, almost like you're scoring the AI's creativity. How do you even define a 3 vs a 4 for that category? Do you have examples or a guide for each number?



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The "below 15, start from scratch" rule is the real gem here. It forces a binary decision and stops the sunk cost fallacy of editing a fundamentally weak draft.

I'd suggest one addition for technical contexts: a **Structure & Scannability** category. Does it use clear headers, bullet points, and data tables appropriately? AI often produces a wall of text, which fails for documentation or API guides. A 1 here means an unbroken paragraph, a 5 means proper hierarchical formatting.

How do you handle scoring when the draft is a mix of great and terrible sections? Do you average it out, or does one catastrophic flaw in Accuracy drag the whole score down regardless of a perfect Tone score?


sub-100ms or bust


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

The accuracy category is a great starting point. We tried something similar for chatbot script drafts, and I'd add a sub-point: check the provided URLs. Ours would occasionally invent support article links that didn't exist. That's an automatic 1.

How do you define a 5 for tone and brand voice? Is it a checklist of keywords, or just a gut feeling from your editor?



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Great catch on the hallucinated URLs. That's a perfect example of an accuracy failure that's easy to miss on a quick read. For us, that's also an instant 1 - it breaks trust completely.

On tone and brand voice, we started with a gut feeling, but that just led to arguments. Now we have a simple checklist that makes a 5 objective:
* Uses our documented "power verbs" (like "configure," "triage," "observe").
* Avoids hyperbolic adjectives ("incredibly," "massively").
* Follows our sentence structure rule: lead with the actionable clause.
* Maintains a calm, diagnostic tone even when describing an incident.

It's not a keyword spam, but if it hits three of those four, it's at least a 4. The final point is the editor's gut check for "does this sound like Sarah from support wrote it?"


Sleep is for the weak


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

That "below 15, start from scratch" rule is a fantastic guardrail. It's the difference between a scoring system and a decision-making system.

For our team, the "Actionability" category was the real eye-opener. We'd get these eloquent, accurate paragraphs that just... ended. No next step, no implied task. Scoring that harshly forced us to be explicit in our prompts, like adding "end with a clear command for the reader." Suddenly a lot of 2s became 4s.


ship early, test often


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

The mix of great and terrible sections is exactly where our team gets stuck too. We've started scoring the entire draft based on its weakest critical section, like a fact error. If you can't trust one part, can you really trust any of it without a full fact-check? That usually means a 1 or 2 in Accuracy, which tanks the total score and triggers the "start over" rule anyway.

I really like the **Structure & Scannability** idea for technical docs. We've had the same issue with AI making a solid point buried in a huge paragraph. Does your team include that category in the total score, making it out of 25 now, or does a bad structure score just serve as a warning to re-prompt?



   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Absolutely love that approach of scoring based on the weakest critical section. That's the method we landed on, too. In email marketing, one wrong link or a misinterpreted segmentation rule invalidates the whole thing, so the whole piece can't score higher than Accuracy. It just saves so much time.

For Structure & Scannability, we rolled it into the main score, bumping the total to 25. But we also found that a low score there (like a 1 or 2 for a wall of text) is often a prompt issue, not a draft-quality issue. So if the total is above 15 but structure is the sole low category, we don't start over. We just re-prompt specifically asking for headers, bullet points, and a clear hierarchy, then re-score the new version.


test everything twice


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

Your point about a low structure score being a prompt issue is a really sharp distinction. We treat it the same way in our data documentation. If a pipeline overview is accurate but a dense paragraph, it's not the LLM's fault, it's ours for not asking for headings.

We took it a step further and made a separate "scaffolding" check. We score the rubric normally, but if Structure & Scannability is below a 3, we mandate a re-prompt with a formatting template before we even consider the other scores for a final decision. This isolates the formatting problem from the content problem, which saves time.

I'm curious, does the "weakest critical section" rule still apply to that new version after a structure re-prompt, or do you reset and judge the new draft fresh?


Extract, transform, trust


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

The separate scaffolding check is smart. Isolating formatting from content is the only way to diagnose a prompt problem versus a content problem.

But I disagree on resetting the judgment. The weakest critical section rule should still apply to the new draft. You re-prompted for structure, but did that re-prompt introduce a new accuracy error or ruin the tone? You have to judge the whole piece fresh, because you changed the inputs. It's a new deliverable.

Otherwise, you're just patching a known flaw and assuming everything else stayed perfect, which it never does.


Your CRM is lying to you.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

You're absolutely right that a re-prompt creates a new deliverable. The core of the weakest section rule is that you must be able to trust the whole piece, and changing the input can break trust in unexpected places.

We learned this the hard way with API documentation. A re-prompt for better structure might cause the model to drop a crucial, but subtle, deprecation note that was in the original wall of text. If we didn't judge the new draft completely fresh, we'd ship a clean-looking but dangerously incomplete spec.

One caveat though: for pure formatting re-prompts that use a strict template (like "reformat this exactly, only adding ## headers"), the risk is lower. But even then, we treat it as a new draft, just with a higher confidence baseline. The rule stays, because the cost of assuming is too high.



   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 2 months ago
Posts: 253
 

You're right about the honest scoring. When a low score means a hard stop, teams don't nitpick over a single point because the outcome is the same: a new draft. It removes the temptation to "bump it to a 3" just to avoid rework.

For our workflow, we score independently first using a shared sheet. Then we huddle to discuss any category where scores differ by more than two points. The independent scoring prevents groupthink, and the discussion focuses only on the big disagreements. It usually takes under ten minutes.

I'm curious, for a team huddle method, do you find people anchor on the first score someone voices out loud? How do you avoid that bias?



   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
Topic starter  

20+ is generous. We scrapped the total score because it lets one great section hide a terrible one. "Weakest link" scoring is the only way that made sense for us. If it hallucinates a feature, the whole thing is a 1. Full stop.

The real problem is treating AI drafts like they're special. You'd redline a junior copywriter for these same flaws. Why give the robot more grace?


CRM is a means, not an end.


   
ReplyQuote
Page 1 / 2