Compare judgements, not totals
The tempting comparison is the overall score: the AI said 72%, the assessor said 74%, close enough. It is also the wrong comparison. Overall scores are sums, and sums hide offsetting disagreement - the AI can be two levels high on one criterion and two levels low on another and the totals will embrace warmly. The unit that matters is the single rubric criterion on a single response: the score the AI recorded against the score the human recorded for the same judgement. That is where agreement or disagreement actually happens, and it is the level at which a rubric gets improved.
Define "reviewed" before you count anything
The single biggest way agreement numbers lie is counting responses nobody looked at. If untouched responses count as agreement, your agreement rate is mostly a measure of how busy your assessors are. A defensible definition: a response counts as reviewed when an assessor engaged with its scores - changed at least one, or explicitly confirmed them. Inside a reviewed response, unchanged scores count as accepted; responses nobody touched are excluded entirely. And release is not review: a release timestamp proves process, not engagement.
The five numbers worth computing
- Exact agreement - the share of reviewed criteria where the human kept the AI score. The headline, and the harshest number.
- Within-tolerance agreement - within one rubric point or level. On a defined level scale, adjacent-level disagreement is usually calibration, not error; this number says whether disagreements are near-misses or gulfs.
- Override rate - the share of reviewed criteria a human changed, with its direction (how often up, how often down). A consistently one-directional override pattern is a rubric or prompt problem wearing a disagreement costume.
- Mean absolute and mean signed difference - how far scores move when they move, and whether the AI runs systematically hard or soft against your assessors.
- Cohen’s kappa (and Pearson r) - chance-corrected agreement, once the sample has enough size and variation to support it. Report it when it is computable; say plainly when it is not.
Sample-size honesty is the whole game
An override rate of 100% sounds damning and an exact agreement of 100% sounds magnificent, and at n=1 they are the same fact wearing different costumes. Below roughly ten reviewed criteria, one decision swings every percentage on the page. The honest presentation at low samples is not a number - it is "not enough reviews yet", with a progress count toward the threshold. Any accuracy claim that arrives without its n attached should be treated as marketing.
Where confidence scores fit
Per-criterion confidence, done properly, measures evidence strength: how directly and unambiguously the answer speaks to the criterion. It is not a quality score - a fully evidenced failure is high confidence, and a polished answer that never addresses the criterion is low confidence. Its correct use is routing: low confidence flags where human review adds the most value, so limited assessor time lands where the AI itself signals uncertainty.
Running the measurement, practically
You do not need a research project. Take one live assessment with a real rubric, have assessors hand-score thirty to fifty responses - blind first, before seeing the AI's draft, for the cleanest read - then compute the five numbers above from the criterion-level pairs. Keep measuring continuously afterwards from routine review activity, so the number is a property of your live operation rather than a one-off benchmark that ages quietly. The measurement itself needs the audit records described in the audit trail guide - without an override ledger there is nothing to compute.
Common questions
- What is the right unit for measuring AI-human agreement in grading?
- The individual rubric criterion on an individual response - not the overall score. Overall scores can agree while every underlying judgement disagrees in offsetting directions. Criterion-level comparison is where disagreement actually lives.
- What agreement metrics should be reported?
- Exact agreement (the human kept the AI score), within-tolerance agreement (within one rubric point or level), override rate (the share of reviewed scores a human changed), mean absolute and signed difference, and a chance-corrected statistic such as Cohen’s kappa once the sample supports it.
- How many reviewed scores are needed before percentages mean anything?
- There is no magic threshold, but below roughly ten reviewed criteria a single override swings every percentage so violently that the numbers are noise. An honest dashboard holds the percentages back at low samples and shows progress toward the threshold instead.
- What do AI confidence scores mean in automated grading?
- A well-designed confidence score expresses evidence strength, not answer quality: how directly the response speaks to the criterion. A clearly evidenced failing answer should be high confidence; an impressive-sounding answer that never addresses the criterion should be low confidence. Confidence is a routing signal for human attention, not a substitute for review.
- Should AI assessment vendors publish accuracy numbers?
- Yes - with the sample size and the method attached, computed from real reviewed data rather than a one-off benchmark. A vendor that will not show its measurement method is asking you to take agreement on faith, which is precisely what assessment is not allowed to do.