Skip to main content
Back to blog
audit trailai assessmentvendor evaluationcompliancebuyer checklist

What Should an Audit Trail for AI-Assisted Assessment Contain? A Buyer Checklist

Adam Broons14 August 20269 min read

When AI is involved in assessment, the audit trail carries more weight than it did before, because there is now a step in the decision that nobody watched happen. A human assessor can be asked what they were thinking. A model cannot. What stands in for that is the record.

This is a checklist you can run against any vendor, including this one. It is written as questions with the answer you should be looking for, because the useful version of this exercise is not reading a features page - it is asking a vendor to show you the record for a single specific result and seeing what is actually in it. Our companion piece covers what an audit trail should contain from the appeals side; this one is about evaluating a product.

The test the whole checklist reduces to

Pick one completed assessment from six months ago. Ask: using only what the system recorded, can you explain to an external reviewer exactly why this person received this outcome on each criterion, including what the AI proposed, what the human changed, and what standard was in force at the time?

If the answer is yes, the trail works. Everything below is a component of being able to say yes.

1. Configuration and rubric versioning

Ask: if I edit a rubric today, what happens to results assessed under the old version?

Look for: the assessed result stays attached to the version of the rubric that was in force when it was assessed, and you can view that old version.

The failure: rubrics stored as mutable records with no history. Six months later the assessment record points at a rubric that has been edited three times, and the criteria you are looking at are not the criteria the learner was judged against. This is quietly one of the most common gaps, and it invalidates the record without anyone noticing. If a vendor cannot show you a prior version of a rubric, they do not have versioning, whatever the page says.

The same applies to anything else that shapes the judgement - grading schema, level definitions, weightings, question text, the prompt or instructions given to the model. If it could change the outcome, the version that applied needs to be recoverable.

2. The override ledger - who changed what, when, and why

Ask: show me a result where the reviewer disagreed with the AI. What does the record contain?

Look for: the AI's proposed score, the final score, who changed it, the timestamp, and ideally a reason. Both values retained, not the original overwritten.

The failure: the system stores only the final score. The AI proposal is discarded on override, so the record shows a human score with no indication that AI was involved at all. This destroys two things at once - the ability to show a regulator that AI assistance was disclosed and supervised, and your own ability to measure how often the tool is wrong.

The override ledger is also the only honest measure of whether human review is real. An organisation with a 0% override rate across a thousand results is not demonstrating an excellent model, it is demonstrating rubber-stamping. You cannot see that without the ledger.

3. Release records - the moment a result becomes real

Ask: at what point can a respondent see their result, and what records that moment?

Look for: an explicit release or sign-off step, separate from the scoring, with a named person and a timestamp. Draft results should not be visible to the respondent.

The failure: results become visible as soon as they are generated. There is no moment where a human took ownership, which means there is no defensible answer to "who decided this". It also means an error caught after generation has already reached the respondent.

The release step is the structural difference between a human being in the loop and a human being near the loop. Ask specifically whether release can be done in bulk without opening individual results - bulk release is a legitimate feature for genuine efficiency, but a system that only offers bulk release is not designing for real review.

4. Model identification

Ask: which model version scored this result, and would I know if you changed it?

Look for: the model and version recorded per result, and a policy of notifying customers when the model changes.

The failure: the model is an implementation detail the vendor swaps silently. This matters more than it sounds. If your cohort scores shift in March and you cannot tell whether the cohort was weaker or the model changed, you have lost the ability to interpret your own data. For anyone doing validation or moderation work, an unrecorded model change contaminates the comparison.

Related question: is the assessment content used to train models? The answer should be no, and it should be in the contract rather than only on a webpage.

5. The submission as received

Ask: is the original response retained, unaltered, or only the scored output?

Look for: the exact submission stored with a timestamp - the actual video, the actual file, the actual text - subject to your retention policy.

The failure: only a transcript or a summary is kept. When a result is challenged, an appeal panel needs to see what the person actually submitted, not the system's rendering of it. Note the tension with data minimisation here: retaining submissions indefinitely is its own problem. The right answer is retention that matches your appeals window and your policy, set deliberately, which we cover in AI assessment and GDPR.

6. Evidence linkage per criterion

Ask: for this criterion score, what in the submission produced it?

Look for: the specific text, timestamp, or file section quoted against each criterion score.

The failure: a criterion breakdown with no linkage. You know the person scored two out of four on criterion three, but nothing connects that to their work. This is the gap that loses appeals, and it is the same gap whether the marker was a human or a model.

7. Export and portability

Ask: can I export the complete record for one result, or one cohort, in a form I can hand to an auditor without your platform?

Look for: a single export containing submission, criteria in force, scores, evidence, AI proposal, overrides, reviewer identity, and timestamps.

The failure: the trail exists but only inside the product, viewable one screen at a time. An auditor asking for forty records should not require forty screenshots. This is also a business-continuity question - if you leave the vendor, does the evidence for past decisions leave with you?

8. Access and change logging on the trail itself

Ask: can an administrator edit a completed assessment record, and would that show?

Look for: either the record is immutable once released, or edits are appended and visible rather than silent.

The failure: a record any admin can quietly change is not evidence. The value of an audit trail rests entirely on the assumption that it was not written after the challenge arrived.

Using this checklist

Send the eight questions to every vendor on your shortlist and ask for a screen share against one real record rather than a written response. The written answers will all sound similar. The demonstration will not.

Watch particularly for the two that most products get wrong: rubric versioning and the retention of the AI proposal after an override. Both are invisible on a features page and both break the record in ways you will not discover until the day you need it.

How Scorafy answers these

In the interest of being held to our own checklist: config and rubric versions are recorded against each result so the standard in force at assessment time is recoverable; overrides retain both the AI-proposed score and the reviewer's decision with identity and timestamp; results require an explicit release before a respondent sees them; the model is identified per result; evidence is cited from the submission for each criterion score; and an Audit Pack export produces the whole record for handing to an auditor. Retention is configurable rather than fixed.

Scorafy is not a certification registry and does not issue credentials - it is the assessment and record layer. More detail on how scoring works and the security overview.

See it live

See AI-powered assessments in action.

Try the interactive demo - no sign-up required.