Skip to main content
Back to blog
ai gradingreliabilityvalidityevidencevendor evaluation

Is AI Grading Reliable? Consistency, Validity, and the Numbers Vendors Quote

Adam Broons14 August 20269 min read

The honest answer is that AI grading is reliable in one specific sense and unproven in another, and most of the confusion in this market comes from vendors quoting the first as though it settles the second.

The first sense is consistency: given the same response and the same rubric, does the tool produce the same judgement every time, and does it apply the same standard to response one and response four hundred. Language models are strong here, considerably stronger than tired human markers. The second sense is validity: is the judgement correct - does a score of "competent" mean what your assessors mean by competent. That question cannot be answered by the vendor at all. It can only be answered against your rubric, your material, and your assessors.

This piece separates the two, explains what evidence citations actually buy you, and gives you a method for testing a tool yourself. We cover the broader capability question in can AI grade open-ended answers; this one is specifically about how far the scores can be trusted.

Consistency is necessary and nowhere near sufficient

A tool that scores the same submission differently on Tuesday than it did on Monday is unusable - you cannot defend a result you cannot reproduce. So consistency is the floor.

But a consistently wrong grader is worse than an inconsistent one, because the error is invisible. Human marker drift is at least noisy and gets caught in moderation. A model applying a subtly wrong reading of your criterion applies it identically across the cohort and nothing looks odd - the failure is systematic, which makes it harder to spot and larger in effect.

This is why "highly consistent" is a weak reassurance on its own. The right question is not whether the tool agrees with itself. It is whether the tool agrees with your assessors on the cases where your assessors are confident, and whether it flags the cases where they would hesitate.

The accuracy percentage problem

You will see numbers like "94% accurate" on vendor sites. Before that means anything you need four things the number is almost never published with.

  • Accurate against what. A ground truth has to come from somewhere. Usually it is a set of human-marked submissions. Whose marking, how many markers, and did those markers agree with each other? If two of your own assessors agree on only 80% of borderline cases, a tool cannot be 94% accurate against "the truth" - there is no single truth to be accurate against.
  • On what material. Accuracy on well-structured essays against a clear rubric will be far higher than on messy workplace evidence against a competency standard. A figure measured on someone else's material tells you nothing about yours.
  • At what granularity. Agreement on the overall pass or fail outcome is a much easier bar than agreement on every criterion at every level. Vendors quoting a single number rarely say which they measured, and the two can differ by twenty points.
  • Adjacent or exact. On a four-level scale, is a score one band away counted as correct? "Within one band" agreement flatters the number substantially. Both are legitimate measures; quoting one without saying which is not.

A number without those four is not a measurement, it is marketing. Treat an unexplained accuracy claim as a signal about the vendor rather than the product - if they will present an unfalsifiable number about their core capability, assume the same standard applies elsewhere.

Scorafy publishes no accuracy percentage. Not because the question does not matter but because the honest answer is that accuracy is a property of a tool plus a rubric plus a set of assessors, and we would rather build the means for a customer to measure it on their own material than quote a number generated on someone else's. There is an accuracy dashboard that compares AI-proposed scores against what assessors actually decided, which is the only version of that figure that is meaningful for a given organisation.

Why evidence citations matter more than the score

Here is the practical shift. A number you cannot check is a claim. A number attached to the specific lines of the submission that produced it is a claim you can verify in about ten seconds.

If a tool says "criterion three: not yet met" and quotes the two sentences where the candidate addressed the criterion, a reviewer reads those two sentences and either agrees or does not. The reviewer is no longer trusting the model, they are checking its work. That is a fundamentally different relationship, and it is what makes AI-assisted marking defensible when a bare score is not.

It also changes the failure mode. A model that scores badly but cites honestly is caught immediately - the quoted evidence does not support the score, and the reviewer sees it. A model that emits scores with no reasoning fails silently until an appeal exposes it. So when comparing tools, weight "shows the evidence for each criterion score" far above any headline accuracy figure: the first is verifiable on your first ten submissions, the second is not verifiable at all.

Confidence flags and where review effort goes

The other thing that separates a serious tool from a score generator is whether it can tell you where it is unsure.

Models are poor at spontaneously expressing doubt - the default is fluent confidence regardless of underlying uncertainty. But a system can flag structurally: the response barely addresses the criterion, the evidence is thin, the case sits near a level boundary, the answer is unusually short or off-topic. Those flags do not require introspection. They come from the shape of the evidence.

Their value is triage. A reviewer with sixty results and three hours does not read all sixty equally - they read the flagged ones properly and spot-check the rest. Per-criterion confidence flags make that allocation rational instead of arbitrary, which is the biggest lever on review quality when time is short.

Human review is a design requirement, not a caveat

For anything consequential - a qualification, a probation decision, a certification - a person has to own the result. That is partly regulatory and partly practical: a decision nobody can explain is a decision nobody can defend when challenged.

The design that works is the model does the reading and proposes scores with cited evidence, a qualified person reviews, adjusts what needs adjusting, and releases. What matters is that the review is real. A reviewer who clicks approve on everything provides no assurance, and the way to tell the difference is whether overrides are recorded - if the override rate is zero across hundreds of results, either the tool is remarkable or nobody is actually looking. Our piece on human-in-the-loop assessment goes further into what a genuine review looks like.

How to test a tool on your own material

This takes an afternoon and is worth more than any vendor claim.

  • Take 20 to 30 already-marked submissions, deliberately including your borderline and unusual cases rather than clean passes and clear fails.
  • Have a second assessor independently re-mark ten of them. This gives you your own human-to-human agreement rate, which is the ceiling any tool can be judged against.
  • Run all of them through the tool with your real rubric, not a simplified version.
  • Compare per criterion, not just on the overall outcome. Note whether disagreements cluster on one criterion - that usually means the criterion is ambiguous rather than the tool being wrong.
  • Read the cited evidence on every disagreement. If the evidence supports the tool's reading, your rubric has a hole in it. That finding alone often justifies the exercise.

Run this against any vendor, including Scorafy. A tool that resists being tested on your own material is telling you something.

In short: AI grading is reliable enough to take the first pass on open-ended work, and not reliable enough to be the final word on anything that matters. Judge tools on whether they show their evidence, flag uncertainty, record what the human changed, and let you measure agreement on your own material. Discount any headline accuracy number that arrives without its method attached.

To see the evidence-and-review pattern in practice, look at how scoring works, or the comparison with using ChatGPT for grading, where the missing pieces are the audit trail and the review step rather than the model itself.

See it live

See AI-powered assessments in action.

Try the interactive demo - no sign-up required.