AI Knowledge Hub

How can organisations measure practical AI capability?

Quick answer

Organisations can measure practical AI capability by defining the decisions and behaviours a person should demonstrate in relevant work, then gathering evidence through realistic tasks, review rubrics, reflection and proportionate workplace application. Completion, knowledge and confidence measures remain useful, but none alone shows that someone can use AI effectively, safely and with sound judgement.

What to remember

Key takeaways

  • Start with observable role behaviour rather than available platform data.
  • Measure task selection, output evaluation and escalation as well as tool operation.
  • Combine controlled task evidence with carefully governed workplace evidence.
  • Interpret results proportionately and state what each measure cannot prove.

Learning platforms make it easy to report who attended, completed a module or passed a quiz. These measures answer useful questions about participation and recall. They do not show whether someone can select an appropriate task, handle information responsibly or challenge a convincing but weak AI output.

Organisations need stronger evidence when the goal is practical capability. Poor measurement can also cause harm if it rewards frequent tool use, encourages people to expose sensitive work or reduces professional judgement to one score.

Measurement should begin with the capability the organisation wants to see.

Capability is a performance question

Practical AI capability is demonstrated through decisions in context. A capable professional can frame a problem, decide whether AI is suitable, choose an approved approach, provide useful context, evaluate the output and recognise when to revise, reject or escalate it.

These behaviours vary by role and task. An underwriter checking a risk summary applies different evidence and consequence criteria from a technical leader reviewing an AI-assisted code change. A universal tool-usage score would ignore the professional work that gives capability its meaning.

Define the target behaviour before choosing a measure. A useful statement might be: “The learner can review an AI-generated summary against the source, identify unsupported claims and record unresolved uncertainty before sharing it.” That statement can guide both practice and assessment.

Common learning measures answer narrower questions

Completion data shows whether someone reached the end of an activity. A knowledge check can test concepts, rules or recognition of a risk. Confidence ratings can reveal how prepared a learner feels and help identify who may need support.

Each source has limits. Completion does not demonstrate application. A quiz may simplify a decision that is ambiguous in work. Confidence can be too low or too high relative to performance. Raw AI usage shows activity but says little about suitability, quality or judgement.

These measures should not be discarded. They become more useful when interpreted as parts of an evidence set. A learner who knows the rules but struggles in a scenario needs a different intervention from someone who performs well but lacks an opportunity to apply the method at work.

Build evidence around a realistic task

Create a task that reflects the intended work without unnecessarily exposing live or sensitive information. Define the source material, permitted tool, intended user, quality criteria, common weaknesses and consequences of error.

Use a rubric that observes the process as well as the final output. Criteria may include whether the learner:

  • Identifies why AI is or is not suitable.
  • Uses permitted information and follows relevant controls.
  • Gives the tool enough context and works iteratively.
  • Traces material claims to evidence.
  • Notices omissions, uncertainty and weak reasoning.
  • Applies role-specific standards.
  • Accepts, revises, rejects or escalates appropriately.

Reviewers need guidance and examples so they apply criteria consistently. Where consequences are higher, involve the relevant domain or risk expertise.

Controlled task evidence can be combined with a later reflection, a manager discussion or a reviewed workplace artefact. A justified decision not to use AI can be valid evidence. Combining sources provides a fuller picture than any single metric.

Measure without distorting behaviour

Measurement changes incentives. If a dashboard rewards the number of AI interactions, people may use a tool where it adds little value. If assessment depends on sharing complete conversations, learners may expose confidential information or feel subject to intrusive monitoring.

Collect only what is necessary for a clear learning purpose. Use controlled scenarios, redacted artefacts or structured reflection where they provide sufficient evidence. Explain who will see the information, how it will be used and how long it will be retained. Local privacy, employment and fairness requirements must shape the design.

Results should be interpreted cautiously. One scenario cannot prove capability across every task. A self-report cannot establish behaviour on its own. Work outcomes may be influenced by tools, processes, colleagues and demand as well as learning.

Reassessment should reflect change and risk rather than an arbitrary universal interval. New tools, policies, responsibilities or recurring errors may justify a fresh check. The purpose is to guide development and responsible practice, not to create a permanent label for the learner.

Example

A financial services customer-operations team completes a fictional case exercise. Reviewers use a rubric covering task suitability, permitted information, source checking, recognition of uncertainty and escalation.

The exercise includes a polished output that omits a material customer detail. Learners must explain what they would change and whether the summary can be used. Several weeks later, each learner provides a short reflection on one approved workplace application or a justified decision not to use AI. Managers confirm the opportunity and support available, without scoring the volume of tool use.

The organisation obtains complementary evidence about knowledge, controlled task performance and workplace application.

FAQs

  • Is confidence a useful AI capability measure?

    Confidence is useful diagnostic evidence, particularly when compared over time or used to identify support needs. It should be calibrated against observed task performance, checking behaviour and recognition of limits because confidence alone can be too low or too high.

  • Should organisations inspect employees' AI conversations?

    Not by default. Inspection requires a clear purpose, proportionality, transparency and compliance with privacy, employment and governance requirements. Controlled scenarios, reviewed artefacts or structured reflections may provide adequate evidence with less intrusion.

  • How often should capability be measured?

    Use a risk- and change-based approach. Reassessment may be appropriate when tools, policies, roles or task consequences change, or when evidence reveals recurring gaps. There is no single interval that fits every role and use.

What's next?

AI Team Based Learning

AI Team Based Learning

AI learning modules designed to develop practical AI capability through short, facilitated modules built around real business scenarios

Our latest learning insights