How Can Risk Teams Test Whether Human Review of AI Decisions Is Meaningful?
Risk teams can test human review by sampling decisions and checking whether reviewers had the information, authority and time to disagree with the AI output. A useful test asks reviewers to explain their own rationale, identify a flawed recommendation and show when they would override or escalate it. Logs of approval alone cannot show that review was meaningful.
Key takeaways
- Define what the reviewer is meant to decide.
- Use sampled cases, including plausible errors, to observe challenge and override.
- Keep a record of the reviewer’s rationale and action.
- Feed weak results into training, workflow and control changes.
A workflow can include a human approval step while still encouraging routine acceptance of AI recommendations. Risk teams need evidence that the reviewer can notice a problem and has a practical route to act. The test should reflect the real decision, its data and its consequences.
What meaningful review requires
Start with the purpose of review. Is the person checking data, weighing a recommendation or deciding whether an outcome is fair? List the evidence available to them and the authority to pause, change or escalate. If these are missing, a training session cannot repair the control by itself.
Build on existing assurance
Traditional quality assurance samples decisions, compares them with source records and looks for patterns of error. Apply the same method to AI-assisted cases. Review a mix of routine and difficult examples, including outliers and cases where the AI answer looks plausible but conflicts with evidence.
Where AI helps
An approved tool may help select samples or group review notes for patterns. It should not score its own reviewers without independent validation. A human assessor should examine whether the reviewer challenged the recommendation for the right reason, not simply whether the final answer matched a model output.
Test and improve the control
Observe a reviewer handling a seeded error, ask for their rationale, and check the override route. Record uncertainty, time pressure and any missing information. If reviewers consistently accept weak outputs, redesign the interface or process as well as training. The ICO’s human-review audit guidance suggests sampling and logging challenge or override; that guidance is under review, so check it again before publication.
Example
A London lender tests a hypothetical AI-assisted affordability review. A control owner gives trained reviewers a case with a credible but incomplete income summary. The reviewers compare it with the application record, record why they disagree and refer the case for a manual decision. Risk staff then examine whether the same response occurs in normal work.
FAQs
-
Is a recorded click enough evidence of review?
No. It shows a step occurred; it does not show the person assessed the evidence or could challenge the recommendation.
-
Should tests include obvious errors?
Include plausible errors and outliers as well as straightforward cases, so the test reflects real judgement.
-
Who should assess the result?
A competent control or assurance owner should assess both the decision and the reviewer’s reasoning.
Get fit for AI
Book a conversation to explore how you can level up your people with the right AI skills.