Skip to content
Notifications
Clear all

Sharing: My team's rubric for scoring AI-generated content quality

23 Posts
22 Users
0 Reactions
4 Views
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Weakest link scoring is the only approach that makes sense for operations content. A monitoring alert runbook with one wrong command is worse than useless, it's dangerous. The total score lets a 5 in "tone" hide that.

We don't huddle. The person who has to *use* the doc or run the procedure scores it. If it fails, they're the ones dealing with the fallout. That enforces brutal honesty.

You're right about not giving it grace. We apply the same rubric to human-written post-mortems. Gets pushback, but it works.


metrics not myths


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

The scoring approach resonates, especially the explicit threshold for starting over. That's where most teams waste time.

I'd push back slightly on "aim for 20+ total" though. In my experience, a total can mask a critical flaw in a single category, like a 1 in Accuracy. We moved to a "minimum floor" for each critical category, particularly for compliance or security comms. A piece can hit a total of 22, but if it scores a 2 on Accuracy, it's dead. The total is useful, but it shouldn't override the weakest critical link.


Review first, buy later.


   
ReplyQuote
(@ellawest)
Estimable Member
Joined: 2 months ago
Posts: 102
 

Completely agree on the total score being a vanity metric that hides failures. You're right to apply a "minimum floor" for critical categories.

We see the same exact problem in identity policy documentation. You can have a beautifully structured, perfectly toned Okta setup guide that gets a total score of 24, but if the "Accuracy" score on one SAML assertion claim is a 2, the entire setup will fail and lock everyone out. The total is meaningless comfort. The 1-5 scale itself is a bit of a trap, honestly. A "2" in Accuracy shouldn't even be a pass; it's a catastrophic fail.

I'd push your point even further: for security or compliance content, the floor should be the *only* score that matters. If any critical category is below a 5, it's a reject. No averaging, no totals. You're not averaging the strength of a chain.


audit logs don't lie


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

You've nailed the core issue with scoring scales. They can imply a false equivalency. A 2 in grammar and a 2 in accuracy are not the same risk category.

Your point about the chain is exactly right, and this is why we map scores directly to an action protocol, not just a number.
* 1-2 in non-critical (e.g., style): Revise.
* 1-2 in critical (e.g., security steps): Hard stop, full re-draft.
* 3-4 in critical: Also a hard stop. We don't accept "mostly accurate" for a SAML claim or a firewall rule. That's where the 5-or-reject floor you mentioned is non-negotiable.

Calling a "2" a catastrophic fail and treating it as such in the workflow is the only way the rubric has teeth.



   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Exactly! Mapping the score directly to an unambiguous action is what makes the rubric operational. It stops the debate over "is this a 2.5 or a 3?" because the action for both is the same: hard stop.

Your note about the false equivalency is huge. We had to split our "Accuracy" category into two for database migration guides: "Factual Correctness" (the command *is* right) and "Procedural Soundness" (the command is *in the right order* for a live system). A 2 in Procedural Soundness means you'll cause downtime, even if every fact is a 5. Different action protocol.

Once you tie a score to a concrete, high-consequence next step, the scoring gets brutally honest, fast.


Backup first.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Oh, that's such a smart way to split it! Factual Correctness vs. Procedural Soundness is brilliant. It makes me think of our webhook implementation guides.

We've been burned by "accurate" payload examples that were technically correct JSON, but were placed in the wrong step of the sequence, so the automation never triggered. The API call is a 5, but the event flow is a 1. Different failure mode entirely, just like your database commands.

Tying the score to the high-consequence next step is what clicked for us too. If a "Procedural Soundness" score is below a 4, our protocol routes it directly to the senior dev who owns that integration. No more passing a "mostly right" draft around.


null


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 2 months ago
Posts: 487
 

The total score is the problem. "Aim for 20+" means a catastrophic 1 in Accuracy can be buried by four 5s in other categories. You're measuring the wrong thing.

We apply the same rubric to post-incident summaries. A 1 in Accuracy there means the report is false. It doesn't matter if the tone is perfect. The whole thing is useless.

Tie each category score directly to an action. A 1-2 in Accuracy? Hard stop, total rewrite. Don't let them average it out.


Five nines? Prove it.


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You're dead on about post-incident summaries. A false RCA is actively harmful because it points future fixes in the wrong direction.

Your point made me think of a nuance: there's a big difference between a "1 in Accuracy" because of a hallucination vs. a "1" because of a critical omission. The latter is arguably worse for a post-mortem. An AI might correctly list the actions taken but completely omit the key log entry that showed the root cause. The facts present are a 5, but the missing fact is a 0, which your rubric flags. That's why the hard stop is so vital.


Keep deploying!


   
ReplyQuote
Page 2 / 2