AI Knowledge Hub

How should a product team design graceful degradation for a critical feature?

Quick answer

Design graceful degradation by deciding which user tasks must still work when a dependency fails, which can pause safely, and what the user should be told. Engineering implements bounded failure behaviour and tests it; Product sets the acceptable user experience and business priority. Never offer a fallback that produces unsafe or misleading results merely to keep a screen available.

What to remember

Key takeaways

  • Classify essential and optional behaviour by user consequence.
  • Define a simple, safe failure mode for each important dependency.
  • Test the degraded path and recovery under realistic conditions.
  • Tell users what remains possible and what will happen to pending work.

A product can appear available while its most important task has stopped working. A slow recommendation service might be tolerable; an unavailable payment authority is different. Teams need a shared answer to what the product should do when part of its environment fails.

Define acceptable degraded service

Graceful degradation means deliberately preserving useful, trustworthy behaviour when part of a system is impaired. List the journeys and dependencies that matter, then classify the consequence of failure. Product decides which journeys must remain accessible and what promise can still be made to users. Engineering identifies the failure signals, technical boundaries and data consistency risks.

Start with failure modes

Established reliability practice uses timeouts, retries within limits, circuit breakers, queues and cached information where appropriate. These mechanisms are useful only when their behaviour matches the task. A stale product recommendation may be acceptable if labelled; a stale account balance may be dangerous. For some failures, stopping the action and explaining why is the safest product behaviour.

Build a controlled response

Design a response for each selected failure: what is shown, whether user input is saved, whether work is queued, and how duplicate requests are prevented. Keep the fallback simpler than the normal path. Product and Engineering agree the user message and any operational workaround. AI can help enumerate scenarios during design, but the team must test them against the real architecture and user consequences.

Test and operate the degraded path

Exercise the degraded path in a controlled environment, including prolonged outage, partial recovery and backlog drain. Monitor the user journey, not only component health. Give operations a clear trigger for entering or leaving degraded mode and a route to communicate status. Review incidents for cases where the fallback hid errors or accumulated work that could not be safely replayed.

Example

Hypothetically, a retailer’s recommendation service fails during checkout. The team keeps basket and payment available and hides recommendations, with monitoring that confirms completed orders still work. Product accepts the temporary loss of suggestions; Engineering tests that retries cannot delay payment. A separate payment-provider outage stops checkout and gives users an honest message rather than recording an unconfirmed order.

FAQs

What's next?

Explore our learning paths

Explore our learning paths

Practical learning to help product and engineering navigate enterprise complexity and deliver exceptional products

Our latest product insights