Skip to main content

Guide - AI assessment assurance

How to measure whether AI grading agrees with human assessors

Every AI grading tool claims accuracy. Almost none of them tell you how they measure it. This guide is the measurement method itself: what to compare, which numbers to compute, and the honesty rules that stop small samples and lazy definitions from manufacturing a headline.

Compare judgements, not totals

The tempting comparison is the overall score: the AI said 72%, the assessor said 74%, close enough. It is also the wrong comparison. Overall scores are sums, and sums hide offsetting disagreement - the AI can be two levels high on one criterion and two levels low on another and the totals will embrace warmly. The unit that matters is the single rubric criterion on a single response: the score the AI recorded against the score the human recorded for the same judgement. That is where agreement or disagreement actually happens, and it is the level at which a rubric gets improved.

Define "reviewed" before you count anything

The single biggest way agreement numbers lie is counting responses nobody looked at. If untouched responses count as agreement, your agreement rate is mostly a measure of how busy your assessors are. A defensible definition: a response counts as reviewed when an assessor engaged with its scores - changed at least one, or explicitly confirmed them. Inside a reviewed response, unchanged scores count as accepted; responses nobody touched are excluded entirely. And release is not review: a release timestamp proves process, not engagement.

The five numbers worth computing

  • Exact agreement - the share of reviewed criteria where the human kept the AI score. The headline, and the harshest number.
  • Within-tolerance agreement - within one rubric point or level. On a defined level scale, adjacent-level disagreement is usually calibration, not error; this number says whether disagreements are near-misses or gulfs.
  • Override rate - the share of reviewed criteria a human changed, with its direction (how often up, how often down). A consistently one-directional override pattern is a rubric or prompt problem wearing a disagreement costume.
  • Mean absolute and mean signed difference - how far scores move when they move, and whether the AI runs systematically hard or soft against your assessors.
  • Cohen’s kappa (and Pearson r) - chance-corrected agreement, once the sample has enough size and variation to support it. Report it when it is computable; say plainly when it is not.

Sample-size honesty is the whole game

An override rate of 100% sounds damning and an exact agreement of 100% sounds magnificent, and at n=1 they are the same fact wearing different costumes. Below roughly ten reviewed criteria, one decision swings every percentage on the page. The honest presentation at low samples is not a number - it is "not enough reviews yet", with a progress count toward the threshold. Any accuracy claim that arrives without its n attached should be treated as marketing.

The same rule applies to the flattering case. A 0% override rate on a handful of reviews is not evidence the AI is right - it may be evidence nobody is really reviewing. Segregating who scores from who releases, and tracking per-assessor calibration, keeps the measurement honest in both directions.

Where confidence scores fit

Per-criterion confidence, done properly, measures evidence strength: how directly and unambiguously the answer speaks to the criterion. It is not a quality score - a fully evidenced failure is high confidence, and a polished answer that never addresses the criterion is low confidence. Its correct use is routing: low confidence flags where human review adds the most value, so limited assessor time lands where the AI itself signals uncertainty.

Running the measurement, practically

You do not need a research project. Take one live assessment with a real rubric, have assessors hand-score thirty to fifty responses - blind first, before seeing the AI's draft, for the cleanest read - then compute the five numbers above from the criterion-level pairs. Keep measuring continuously afterwards from routine review activity, so the number is a property of your live operation rather than a one-off benchmark that ages quietly. The measurement itself needs the audit records described in the audit trail guide - without an override ledger there is nothing to compute.

Common questions

What is the right unit for measuring AI-human agreement in grading?
The individual rubric criterion on an individual response - not the overall score. Overall scores can agree while every underlying judgement disagrees in offsetting directions. Criterion-level comparison is where disagreement actually lives.
What agreement metrics should be reported?
Exact agreement (the human kept the AI score), within-tolerance agreement (within one rubric point or level), override rate (the share of reviewed scores a human changed), mean absolute and signed difference, and a chance-corrected statistic such as Cohen’s kappa once the sample supports it.
How many reviewed scores are needed before percentages mean anything?
There is no magic threshold, but below roughly ten reviewed criteria a single override swings every percentage so violently that the numbers are noise. An honest dashboard holds the percentages back at low samples and shows progress toward the threshold instead.
What do AI confidence scores mean in automated grading?
A well-designed confidence score expresses evidence strength, not answer quality: how directly the response speaks to the criterion. A clearly evidenced failing answer should be high confidence; an impressive-sounding answer that never addresses the criterion should be low confidence. Confidence is a routing signal for human attention, not a substitute for review.
Should AI assessment vendors publish accuracy numbers?
Yes - with the sample size and the method attached, computed from real reviewed data rather than a one-off benchmark. A vendor that will not show its measurement method is asking you to take agreement on faith, which is precisely what assessment is not allowed to do.

This method runs live in Scorafy

The Accuracy page computes every metric on this guide from your organisation’s own reviewed responses - and holds the percentages back until the sample can carry them.