AI Knowledge Hub

How should a Product Engineering team choose its first AI use case?

Quick answer

A Product Engineering team should choose a first AI use case that addresses real work, is valuable enough to learn from and is bounded enough to verify and reverse. The best starting point has an accountable owner, representative demand, suitable data and reliable feedback. It is rarely the most spectacular opportunity or merely the easiest demonstration.

What to remember

Key takeaways

  • Select a delivery problem, not a feature to showcase.
  • Balance useful value with bounded risk and strong feedback.
  • Use representative work so the result informs a real decision.
  • Prefer the simplest AI capability that can test the hypothesis.

The first use case shapes what an organisation learns about AI-enabled Product Engineering. A trivial demonstration may create enthusiasm but reveal little about delivery. An ambitious customer-critical project may expose the team to avoidable risk before its controls and working practices have been tested.

Selection is therefore a Product and Engineering decision. It should connect an actual constraint or opportunity with work that produces credible evidence.

Start with real demand

Identify recurring work, queues, delays or quality problems in the current value stream. Ask where Product or Engineering effort is consumed, where feedback is slow and which work is sufficiently understood to describe. Candidate uses might include test preparation, routine maintenance, codebase explanation, documentation or bounded feature changes.

State a hypothesis in operational terms. For example: giving the team AI support for routine dependency analysis may reduce investigation effort without increasing rework or security exceptions. Avoid goals such as “increase AI adoption” or “generate more code”, which measure activity rather than an improved outcome.

Confirm that the problem matters to the team and has an owner. A centrally selected pilot that competes with urgent delivery may receive little serious use, making low adoption impossible to interpret.

Score value, learnability and exposure

Compare candidates across three groups of factors. Value includes frequency, delay, human effort and relevance to product outcomes. Learnability includes clear acceptance criteria, representative examples, a baseline and fast feedback. Exposure includes data sensitivity, system consequence, architectural reach, reversibility, permissions and the strength of current controls.

The preferred first use case usually sits in the useful middle: meaningful enough that improvement matters, but bounded enough that errors can be detected and contained. Low-risk does not have to mean low-value. A common maintenance queue can be commercially relevant without involving production authority or a critical redesign.

Record why candidates were rejected or deferred. A high-value use may become suitable after the team improves tests, creates a sandbox or clarifies data policy.

Match the AI role to the work

Choose the least agency needed to test the hypothesis. An assistant may help a person analyse or draft within an existing task. An agent may inspect a repository, change files and respond to test feedback. Greater autonomy introduces more permissions, monitoring and recovery needs and should earn its place through expected learning value.

Consider whether AI is necessary. Improving requirements, removing a dependency, automating a deterministic check or fixing the delivery environment may address the constraint more directly. The purpose of selection is to improve capacity or outcomes, not to ensure AI wins every comparison.

Time-stamp assumptions about tool capability. Test candidate tools against the organisation’s languages, repositories and work rather than relying on a demonstration or general benchmark.

Define the decision the use case will support

Before beginning, state what will happen after the trial. The choices may be to stop, repeat with changes, adopt within the current boundary or evaluate a broader use. Define minimum evidence for that decision across quality, lead time, human effort, rework, cost, experience, security and operational effects.

Include a credible comparison. This may be a recent baseline, alternating similar tasks or a small contemporaneous control, depending on available demand. Do not overstate what a small pilot can prove. Its purpose is to reduce uncertainty in context, not establish a universal productivity claim.

Choose enough representative work to expose normal variation, but keep the scope small enough to supervise closely. A good first use case leaves the team with reusable knowledge about its work, environment and controls even if the AI approach is not adopted.

Example

A platform team considers three candidates: generating a showcase application, diagnosing recurring build failures and changing production access policies. The showcase has little connection to actual demand, while access policy changes have a high consequence and weak recovery evidence.

The team chooses build-failure diagnosis because the queue is real, historical examples exist and engineers can verify proposed causes without allowing the assistant to alter production. It compares investigation effort, resolution quality and repeated failures with the current approach. The trial can therefore inform a practical adoption decision rather than only demonstrate the tool.

FAQs

What's next?

Our latest product insights