Skip to content
Notifications
Clear all

Showcase: Built a custom template for our engineering retros, screenshot inside

11 Posts
11 Users
0 Reactions
29 Views
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
Topic starter   [#21610]

Our engineering retros had become a predictable, low-value ritual. We were cycling through the same three vague action items while our post-incident review cadence slowed to a crawl. The standard "What went well? What didn't? What can we improve?" template in Fellow was failing us; it produced qualitative fluff that was impossible to trend or act upon.

I built a custom template to enforce a data-driven, SRE-aligned retrospective process. The goal is to shift discussions from subjective feelings to measurable system behavior and concrete toil reduction. The template forces quantification and ties discussions directly to our service level objectives.

**Core Sections of the Template:**

* **Incident Context (Pre-Filled):** This pulls in the incident ID, timeline, and impacted services from our PagerDuty integration. No more starting from scratch.
* **Impact Quantification:** A table requiring numeric inputs for:
* Duration of user-impacting degradation (in minutes)
* Scope of impact (percentage of user base or request volume)
* SLO/SLI burn (e.g., "Error Budget consumed: 2.1 hours")
* **Root Cause Analysis (Structured):** A multiple-choice dropdown for primary cause category (Code Defect, Configuration Change, Platform Failure, Dependency Failure, Unknown) followed by a free-text field for the precise trigger. This allows us to generate Pareto charts of failure modes over time.
* **Action Items with Accountability:** Each action item must be tagged with:
* Type: `Mitigation` (fix the immediate issue), `Prevention` (prevent this exact failure), or `Optimization` (reduce toil in the response).
* Owner: Assigned directly in Fellow.
* ETA: A concrete date, not "next sprint."
* **Follow-Up Tracking:** A dedicated field for the link to the post-mortem document in our wiki and the ticket number for the primary prevention work.

Here is the JSON structure of the template as exported from Fellow's API, which you can adapt. The key is in the `items` array, where each object defines a section or field with specific `attributes` to constrain input.

```json
{
"template": {
"name": "SRE Incident Retrospective",
"items": [
{
"type": "text",
"attributes": {
"label": "Incident ID & Link",
"description": "From PagerDuty or incident management system.",
"required": true
}
},
{
"type": "table",
"attributes": {
"label": "Impact Quantification",
"columns": [
{"heading": "Metric", "required": true},
{"heading": "Value", "required": true},
{"heading": "Method of Calculation", "required": false}
],
"description": "Quantifiable impact only. Use SLI/SLO data."
}
},
{
"type": "dropdown",
"attributes": {
"label": "Primary Cause Category",
"options": ["Code Defect", "Configuration Change", "Platform Failure", "Dependency Failure", "Unknown"],
"required": true,
"multiple": false
}
}
]
}
}
```

The results after three months are telling. We've reduced the average time to publish a finalized retro from 5 days to 1.5 days. More importantly, 80% of our action items are now of the `Prevention` type, compared to roughly 30% before. The structured data allows us to feed metrics into our FinOps dashboard, drawing a direct line between incident causes and cloud cost drivers (e.g., a specific platform failure mode correlated with auto-scaling overprovisioning).

The template isn't a silver bullet—it requires discipline to fill in correctly—but it has transformed our retros from conversational post-mortems into actionable engineering audits. I'm interested to see if others have taken a similar structured approach or if you've found different key metrics to be indispensable for continuous improvement.

-- alex



   
Quote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Quantifying impact is good, but I've seen this slide into another flavor of box-ticking.

Your table requires "Duration of user-impacting degradation". Great. But who decides that start and stop timestamp? Is it the first alert? The first user report? When the graph looked weird? Without a strict, shared definition, those minutes are just political numbers.

Same with "Scope of impact (percentage of user base)". Is that a guess? An extrapolation from sampled logs? Feels like you're just swapping vague qualitative fluff for vague quantitative fluff.

Hope you've got a runbook for *measuring* the metrics before you demand them in the retro.


-- old school


   
ReplyQuote
(@claraj)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Exactly. The numbers become arbitrary inputs for some "objective" dashboard, which then becomes the new gospel. I've watched teams burn more cycles debating whether an incident was 47 or 52 minutes than fixing the root cause. You're not measuring system behavior, you're measuring who's better at negotiating the narrative.


Prove it


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Yeah, that's a scary side effect I hadn't considered. Turning the retro into a negotiation over metrics defeats the whole purpose.

So is the fix to have those definitions locked in *before* an incident happens, like in a runbook? Or is this just an unavoidable trap with any data-driven approach?


Still learning.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The trap isn't data. It's missing definitions. You have to decide the rules before you play the game.

If your start time is "first alert," then you write a rule saying exactly that. Define the alert. Document the source system. Now it's not a negotiation, it's a lookup. Same for impact scope. Use your real user monitoring platform's calculated percentage.

If you can't define the rule, then you can't measure the metric. And if you can't measure it, you have no business demanding it in a retro template.


Beep boop. Show me the data.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

You've put your finger on the core challenge of moving from a qualitative to a quantitative retro. "Political numbers" is exactly the risk when definitions aren't crystal clear.

This is why the template must be paired with a ratified operational playbook. If you don't have an agreed source of truth for an incident's start time, the template field shouldn't be a blank box. It should be a direct link to the immutable alert from your monitoring system. The goal isn't to ask for a number, it's to force reliance on a predefined system of record.

Without that lockstep, you're right. You're just trading one form of vagueness for another, and arguably a more damaging one because it wears the mask of objectivity.


—daniel


   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Exactly. The template shouldn't even render if the linked data source isn't populated. Otherwise you're just creating a new manual step.

Our SLO error budget tool exposes incident duration via API. That's the only field our retro template accepts for the timing question. No typing, no debate.

If your monitoring stack can't provide the metric, the template field should fail closed. Forces you to fix the instrumentation first.


Ship it, but test it first


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

You've nailed the primary failure mode. It's the transformation of a retro into a metric litigation session, where the most polished arguer wins.

But I think the deeper issue is that even with locked definitions, you're still optimizing for the dashboard, not the system. The team's focus shifts to managing the official record, asking "how do we *classify* this to avoid burning budget?" instead of "how do we *prevent* this?"

It becomes performance theater with numbers.


Show me the TCO.


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You're onto something I've felt but never put into words. The "optimizing for the dashboard" trap is real, and it creeps in quietly.

We fixed this by adding one rule to our retro template: every action item must include a manual validation step that can't be automated. Something like "a human from support confirms the new error message is clearer." It forces the discussion back to real-world impact, not just moving a metric.

If your only goal is to make the dashboard numbers look good, you'll end up playing the numbers game. The template needs to guard against its own success.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@gabrielm)
Reputable Member
Joined: 2 months ago
Posts: 253
 

This is a smart pivot away from vague action items. I'm curious, since you're moving away from the standard Fellow template, what made you choose to build this custom solution directly instead of using a more structured platform like Jira? I'm trying to understand the trade-offs between a custom template and something like a dedicated post-mortem workflow in Jira Service Management, which also tries to enforce similar data fields.



   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 6 months ago
Posts: 427
 

Nice! Pulling incident context from PagerDuty is a huge time saver. I've been trying to do something similar by linking my Grafana alert state changes to our incident log. It's a lot harder than it sounds, though.

How are you handling false positives? Like, if a page fires but gets resolved in two minutes before anyone joins, does that still pre-fill the context in the retro template? Or is there a filter for that?



   
ReplyQuote