What makes a retrospective useful
The most common retrospective failure is a document that got filled out and never read. Every box has something in it, the action items are filed, and six months later the same incident happens again. Nobody points to that retrospective and says it changed how the team works. That document is a receipt: proof the process ran, with no learning attached. Here’s a useful test. Give the retrospective to an engineer who joins your team six months from now. Would they understand:- What happened?
- Why it made sense to the people involved at the time?
- What the organization now knows that it didn’t know before?
Retrospective vs. postmortem: is there a difference?
In practice, the terms describe the same activity: a post-incident review. “Postmortem” is the older, more common term. “Retrospective” has gained ground because it avoids the morbid framing and emphasizes learning over autopsy. Some teams draw a soft distinction, using postmortem for the written document and retrospective for the meeting and process around it, but there is no industry-standard difference. Pick one term, define it, and use it consistently. For more on the history of both terms, see retrospective vs. postmortem.How to run an incident retrospective
These seven steps take a retrospective from scheduling to a shared set of learnings.1
Schedule It Quickly
Hold the retrospective within a few business days of resolution, ideally within a week. Memory decays fast, and the details that matter most (what people saw, what they believed, why they acted) are the first to go. Invite every responder who was involved. Managers can attend, but they shouldn’t stand in for the people who were there.
2
Build the Record Automatically
Reconstruct what happened in order: when the issue started, when it was detected, key decisions, mitigation attempts, and resolution. Pull from your incident channel, monitoring alerts, deploy logs, and the incident timeline. Nobody should ever have to type a timestamp: every minute an engineer spends copying metadata is a minute they aren’t thinking. A shared, factual record settles disagreements about “what happened” before anyone discusses “why.”
3
Write the Narrative From What People Knew
Document the incident from the responders’ point of view, including the wrong theories and dead ends. The dead ends are where the learning is: if the team spent 40 minutes on the database when the problem was DNS, that tells you something important about your observability and your mental models. For each key decision, ask “What made that decision seem reasonable in the moment?” Hindsight makes every choice look obvious; this section’s job is to undo that. Avoid counterfactuals (“they should have checked the dashboard”) and name systems, not people, as points of failure.
4
Look for Contributing Factors
Real incidents almost never have one root cause. Several things had to line up: a config change, a missing alert, an out-of-date runbook, a Friday deploy, an on-call engineer three weeks into the job. Take any one away and maybe there’s no incident. Asking “why” five times usually stops at a person; contributing factors keep you looking at the system. Keep asking “What else had to be true?” until the answers stop being technical and start being organizational: staffing, priorities, deadlines, who knew what. Those are usually the factors that matter most. Slow detection and slow mitigation are findings just as much as the triggering defect.
5
Have the Conversation
The meeting is where the learning happens. People reconstruct what happened and argue about why, and that struggle is what changes how they think. A retrospective published an hour after resolution with no discussion has automated away the one step that mattered. Only the person who made a decision can explain why it made sense, so ask them.
6
Agree on a Few Follow-Up Actions
Turn findings into concrete follow-ups: fix the bug, add the missing alert, update the runbook, add a guardrail to the deploy pipeline. Every action needs a single owner, a due date, and a priority. Fewer is better: five actions that get done beat fifteen that never leave the backlog. And not every learning has to become a ticket. Sometimes the output is simply that three teams now understand the payment service better. That still counts.
7
Share the Learnings
Publish the retrospective where the whole engineering organization can read it, and announce it in a team meeting, a newsletter, or a dedicated channel. Other teams likely share the same failure modes. An unread retrospective only teaches the people who were already there.
Design a template that teaches
Your template decides whether you get a retrospective or a receipt, because it shapes what people think about. A field labeled “Root Cause” gets you one sentence, sometimes with a person’s name in it. A field that asks “What conditions had to be true at the same time for this to happen?” gets you something useful. The Best Practice template, the default for new Rootly accounts, is built around seven sections. Each one exists for a reason, and each can be pushed a little further with the right question in its help text.
A few notes on why the template is shaped this way:
- Impact must be honest. If you catch yourself writing “some users may have experienced degraded performance,” stop and say what it was: “4,200 customers couldn’t log in for 38 minutes.” Honest impact is what keeps leadership trusting the process. And if you don’t have the hard numbers, that’s a sign you may need better metrics.
- What went well is a first-class question. Your ability to respond is something you can learn and strengthen, just like your failures. Somebody on the call almost certainly did something smart that isn’t written down anywhere. The answer to “whose expertise saved us?” is both something to reinforce and a risk you didn’t know about.
- The timeline goes last. If you open with 60 timestamps, the thinking gets buried. Rootly builds the timeline for you, and the Curated Timeline AI block trims it to the key moments, so put it in the appendix.
Write help text as questions, not labels
Help text is the part of the template that does the teaching. “Describe the impact” gets a sentence. “Who was affected, how would they describe it, and what did it cost them?” gets a real answer. When you customize a template, rewrite every section’s guidance as a question.Add a block for what you still don’t know
Consider adding one custom section such as Open Questions: what we still can’t explain. A retrospective that claims to understand everything is usually wrong. Writing down what’s unresolved is more honest, and it’s often where the next incident will come from.Match the effort to what’s at stake
One template for every incident means your SEV0s don’t get enough attention and your SEV3s get too much process. Near misses are cheap lessons, so don’t skip them, but don’t make them expensive either. Think in three tiers:
In Rootly, retrospective processes right-size the follow-up steps by severity, team, or incident type, and retrospective workflows can apply a different template to each. Liquid conditions in a template can also show or hide sections, such as an executive summary that only appears for SEV0 and SEV1. For the heavy tier, start from the Major Incident starter template, which favors completeness and causal depth over brevity. Match the effort to the stakes and people won’t resent the process.
Use AI for recall, not meaning
The learning is what a retrospective produces, and it happens in people’s heads, through the struggle of reconstructing what happened and arguing about why. If AI does that struggling for you, you end up with a polished document and an organization that learned nothing. The simplest rule: AI is excellent at recall and should never be trusted with meaning.Where AI earns its place
- Metadata and timeline curation. Pulling key moments out of a 400-message Slack channel and a two-hour bridge call is grunt work people are bad at.
- A first draft of What Happened, as a memory aid. Who said what, and when, including things responders have already forgotten.
- Catching dropped action items. The Follow-Up Actions block suggests actions that were promised during the incident but never logged.
- Answering questions about the incident. “When did the error rate first climb?” “Who approved the rollback?”
- Checking blameless language. The Best Practice template’s blameless instruction set already tells every AI block to use roles instead of names and to separate confirmed from suspected.
Where people stay in charge
- Deciding what the contributing factors and learnings mean. AI only sees what was written down, not what was said in DMs, in the hallway, or in the on-call engineer’s head at 3 a.m. The most important factors are often the ones nobody typed.
- Deciding priorities. Which follow-up matters most depends on your roadmap, your risk appetite, and your organization.
- Holding the meeting. The conversation is where the learning happens, and no draft replaces it.
Set up AI so the draft opens the conversation
The AI variant of the Best Practice template drafts Contributing Factors and Learnings along with the other narrative sections. Treat those drafts as the starting point for the meeting, not its outcome. In Rootly:- Tell AI to ask, not answer. Each AI block takes custom instructions on top of the template’s blameless rules. For Contributing Factors, try: “List possible contributing factors and cite the source for each. Don’t rank them and don’t draw conclusions.” Or add a custom block: “List the questions the retrospective meeting needs to answer.” Now the AI sets the meeting agenda instead of writing the outcome.
- Check the sources. Every AI block shows how it was generated, including the sources and the prompt behind it. If a claim doesn’t trace back to something real, delete it.
- Nobody publishes an AI block they haven’t argued with. Make it a team norm. The draft is a proposal to push back on, not an answer to accept at face value.
- Rate the output. Use the thumbs up and down on each block; that feedback goes into how Rootly evaluates output quality.
Common mistakes
- Treating a completed document as a finished retrospective. Every box filled in doesn’t mean anyone learned anything. Apply the new-engineer test.
- One process for every incident. SEV0s get too little attention and SEV3s get too much. Tier your retrospectives.
- Blame in disguise. “Human error” as a root cause, or a narrative written as an accusation. If a person “caused” the incident, the system that let one mistake cause an outage is the real finding.
- Root-cause tunnel vision. Stopping at the first plausible technical cause and missing the detection, response, and organizational factors around it.
- Letting AI write the conclusions. AI-drafted contributing factors and learnings reflect only what was written down and miss tone, DMs, and organizational dynamics.
- Too many action items with no follow-through. Unowned, undated items quietly expire. Keep the list short and review open items regularly.
- Waiting weeks to run it. Stale memories produce vague timelines and generic conclusions.
- Writing it and telling no one. The document is a means; the learning is the point.
Run retrospectives in Rootly
Most retrospective failures are process failures: steps forgotten, documents never started, action items never tracked or followed up on. Rootly automates that scaffolding so your team can spend its time on the thinking.- Start from the Best Practice template. Everything this guide recommends (contributing factors over a single root cause, blameless narratives, what went well as a first-class question, the timeline as an appendix) is encoded in it. New accounts get it by default, and existing customers can select it from the template list. Start there, then add your own questions.
- Define retrospective processes. Set ordered steps (gather data, write the document, host the review, create action items, share the report) with due dates, role-based assignees, and reminders, and right-size them by severity, team, or incident type.
- Automate with retrospective workflows. Trigger on retrospective lifecycle events, for example to create the doc in Confluence or Google Docs, notify leadership channels, and open follow-up tickets when a retrospective is published.
- Track action items. Give every follow-up an owner, a due date, and a status, with exports and dashboards so open items stay visible long after the incident channel goes quiet.
- Draft with Rootly AI in retrospectives. AI blocks draft sections from the incident’s timeline, Slack discussion, and bridge-call transcripts captured by Meeting Scribe. Ask Rootly AI answers questions about the incident in a chat that’s private to you, so bring anything useful back into the retrospective. Reviewers still edit and own the narrative: the AI removes the busywork, not the judgment.