AI Knowledge Hub

How should Product and Engineering learn from a production incident?

Quick answer

Learn from a production incident by recording what users experienced, reconstructing events from evidence, examining the conditions that allowed the failure and assigning a small set of verifiable improvements. Product assesses customer impact, communication and priority; Engineering examines technical causes and controls. Review whether actions worked, and share lessons without blaming individuals for decisions made with limited information.

What to remember

Key takeaways

  • Separate user impact, timeline, contributing conditions and actions.
  • Use logs and firsthand accounts to test explanations.
  • Prioritise changes that reduce likelihood or impact and give each an owner.
  • Check effectiveness later; a completed document alone does not improve reliability.

After service is restored, teams often rush back to planned work. The immediate fix may address a symptom while the same conditions remain. A structured review helps Product and Engineering understand what customers experienced and decide which changes are worth making.

Decide when and how to review

Set review triggers before an incident, such as significant user disruption, data loss, recovery difficulty or a surprising near miss. Scale the effort to the event and its learning potential. Preserve key logs, decisions, support contacts and customer communications. Product describes affected journeys and obligations to users; Engineering records system behaviour, mitigation and remaining uncertainty. Involve operations and support where they saw effects the technical team could not.

Reconstruct the event with evidence

Build a timeline of what was known, what actions were taken and when the user impact changed. Distinguish observed facts from hypotheses and avoid filling gaps with confident stories. Ask why the system, alerts, runbooks and team context made each action reasonable at the time. A blameless review examines conditions that shaped behaviour; it still names accountable owners for improvements. An urgent legal or security investigation may need its own specialist process.

Choose improvements that matter

Group candidate actions into prevention, detection, mitigation and recovery. Prefer specific, testable actions over vague instructions to be more careful. Product and Engineering weigh user consequence and capacity before committing. A small improvement to monitoring or rollback may be more useful than a large redesign, depending on the evidence. AI can summarise a reviewed timeline or search records, but human participants must verify sensitive details and causal claims.

Follow through and share learning

Assign an owner and due review point to each accepted action, then check whether the change actually works through a test, drill or later operational evidence. Update support guidance and user communication where appropriate. Share findings with teams that operate similar dependencies. Revisit assumptions if a later incident contradicts the explanation. The review should improve the product and delivery system, rather than serve as a formality.

Example

Hypothetically, a booking service starts returning errors after an upstream provider slows down. Once restored, Product gathers affected journey and support evidence. Engineering reconstructs request traces and identifies unbounded retries that amplified the delay. The team tests bounded retries, updates the failure message and checks the mitigation in a rehearsal. The review records remaining uncertainty about provider behaviour.

FAQs

What's next?

Explore our learning paths

Explore our learning paths

Practical learning to help product and engineering navigate enterprise complexity and deliver exceptional products

Our latest product insights