AI Knowledge Hub

How Often Should AI Models Be Reviewed and Revalidated?

Quick answer

There is no universal review interval. Monitor live performance and control outcomes continuously at a cadence proportionate to use and risk, schedule broader reviews according to impact and rate of change, and revalidate when a material trigger occurs. Triggers include changed source formats or business mix, model or prompt updates, new rules or reference data, rising exceptions or overrides, incidents and expanded use.

What to remember

Key takeaways

  • Ongoing monitoring, periodic review and revalidation are different controls.
  • Set cadence from impact, change rate and ability to detect harm.
  • Revalidate the components affected by a material trigger.
  • Record evidence, approval, remediation and any restricted use.

An AI-supported bordereaux workflow is approved for a defined use under defined conditions. Those conditions can change as coverholders, formats, business mix, reporting requirements and technology evolve.

A fixed annual or six-monthly date cannot provide all the necessary assurance. A problem may emerge immediately after the scheduled review, while a stable low-impact component may not justify the same review intensity as a material and rapidly changing use.

A stronger control model combines live monitoring, planned review and event-triggered revalidation.

One calendar interval cannot represent every risk

Review frequency should follow intended use, impact, complexity, rate of change and the organisation's ability to detect a problem. An AI suggestion used only to prioritise low-materiality work creates a different exposure from output that may be accepted into finance, claims or regulatory reporting.

The term model review can also hide the real scope. A production workflow may include a model, prompt, mapping configuration, deterministic rules, reference data, target schema, interfaces and human controls. Any of these can change the outcome.

A model may remain unchanged while a coverholder alters a workbook or the organisation expands into a new class. Conversely, a supplier may update a model while the input format appears stable.

The review plan should inventory these components, versions, owners and dependencies. It can then set different monitoring and assurance activities without pretending one date covers the whole service.

Established model and control assurance is risk-based

Traditional assurance begins with a defined purpose, documented assumptions, testing, approval, named ownership and independent challenge proportionate to risk. These principles remain useful for AI-supported workflows.

The PRA's current model-risk principles apply to specified banks rather than serving as a direct rule for every delegated authority insurance workflow. They nevertheless illustrate an important general discipline: validation intensity and frequency should be determined by the firm, model and risk rather than a universal minimum interval.

DA organisations should apply the legal, regulatory and governance requirements relevant to their own status, jurisdiction and use. The operational design still needs a clear distinction between the owner who monitors performance, the users who observe outcomes, the validators who challenge evidence and the authority that approves continued use.

Periodic review considers whether the purpose, assumptions, risk classification, controls, evidence and ownership remain suitable. Independent input helps prevent the team that built or operates the workflow from marking its own work without sufficient challenge.

AI workflows need live outcome monitoring

Monitoring should run often enough to detect deterioration before the next formal review. The cadence may differ by signal: service health can be observed continuously, while validated accuracy may depend on outcomes collected over a meaningful period.

Track input structure and quality, relevant data segments, mapping and validation performance, confidence where meaningful, exceptions, reviewer corrections, overrides, reconciliation breaks, incidents and queue ageing. Compare them with approved baselines and limits.

Monitor versions of models, prompts, rules, mappings and reference data so that a change in outcome can be connected to a change in the service. Supplier updates should enter the same impact and release process as internal changes.

An alert starts investigation; it does not prove that the model has drifted. Increased exceptions may result from a changed coverholder format, a new business mix, a broken interface or a stricter rule. Diagnosis determines what needs revalidation or remediation.

Material change triggers targeted revalidation

Do not wait for the next scheduled review after a material trigger. Relevant events include a new coverholder or class, changed source structure, expanded downstream use, model or prompt release, altered mapping or rule, new reference data, sustained performance movement, rising corrections, an incident or a control failure.

Assess which components and populations are affected. A new mapping rule may need focused regression and end-to-end testing rather than redevelopment of the model. A major model update or new use may require broader revalidation.

The evidence should cover representative current data, material segments, known difficult cases, human-review outcomes, limitations and control effectiveness. Record the scope, tests, findings, reviewer independence, decision and any conditions.

Possible outcomes include continued use, restricted use, stronger review, remediation, rollback, temporary fallback or withdrawal. Track actions to closure.

The resulting cadence is explicit but not arbitrary: ongoing signals, scheduled governance and material events work together to keep the workflow within its approved purpose.

Example

A hypothetical managing agent uses AI-assisted field mapping for stable property bordereaux. Live monitoring shows consistent mapping outcomes and reviewer corrections within the approved range.

The firm then onboards a coverholder writing a new specialty class. Mapping overrides rise for several class-specific fields. The new use and override pattern trigger immediate targeted revalidation rather than waiting for a calendar review.

The validation lead tests the affected segment, identifies missing reference examples and recommends mandatory review until remediation passes regression and end-to-end checks. The existing stable property segment continues under its approved monitoring plan.

FAQs

What's next?

Talk us through your DA process

Talk us through your DA process

Book a conversation to explore where AI could help improve delegated authority data flow, validation and operational control.

Our latest insurance insights