Most onboarding evaluation measures the wrong thing. Ask a new starter at day 30 whether onboarding was useful and you learn about the experience. Give them a multiple-choice test on the policy handbook and you learn whether they can recognise a correct answer. Neither tells a manager what they actually need to know: can this person now do the work the role requires, and where are the gaps.
Evaluating onboarding properly means treating it as a competency question rather than a satisfaction survey or a recall test. Here is what to measure, how to shape it across 90 days, how to write the rubrics, and where AI evaluation helps - including where it should stay away.
Three different things get called "onboarding evaluation"
Kirkpatrick's four levels are still the clearest way to separate them. Level one is reaction - did people find it useful. Level two is learning - did they acquire the knowledge and skill. Level three is behaviour - are they doing the work differently on the job. Level four is results - did the business outcome move.
Almost every onboarding programme measures level one, because a survey is easy to send. A few measure level two with a quiz. Very few measure level three, which is the level that predicts whether the new starter succeeds - and the gap is not laziness, it is that level three needs someone to look at real work and judge it. You do not need all four. You need level three, sampled well, plus enough of level two to know whether a knowledge gap or a judgement gap is causing the problem.
Start from the tasks, not the content
The most common design error is building the assessment around what was taught, which produces questions about the induction deck rather than the job. Instead, list the six to ten things this person will be doing by month three, then work backwards to what evidence would show they can do each one.
For a support role: triage an ambiguous ticket, explain a technical problem to a customer, decide when to escalate. For a care role: complete a handover, identify a risk in a scenario and state the mitigation, document an incident to standard. For sales: qualify a lead against your criteria, handle an objection, write a follow-up a manager would be happy to send.
Every one of those is open-ended. There is no key. That is the point - the tasks that matter in a job rarely have a single correct answer, which is why quiz-based onboarding evaluation feels hollow.
The 30-60-90 shape
Spread the evidence across three checkpoints, with the standard rising each time.
- Day 30 - orientation and safe execution. Can they find the right information, follow the process, and recognise when something is outside their remit. Evidence: a short written walkthrough of a routine case, plus where they would look things up. Low stakes and diagnostic - the output is a list of what to reinforce, not a score anyone acts on.
- Day 60 - core tasks with light supervision. Can they perform the main tasks of the role to an acceptable standard without step-by-step direction. Evidence: real or realistic work samples - an actual ticket response, a drafted plan, a recorded explanation. This is where judgement starts to show and where genuine gaps appear.
- Day 90 - independent performance and decision quality. Can they handle an ambiguous case and justify the call they made. Evidence: a scenario with no obvious right answer, assessed on the reasoning as much as the outcome. If this checkpoint has consequences attached - probation, role confirmation - treat it as a consequential decision and design it accordingly, which means a named human owns the result.
The shape matters more than the exact dates. You are building a trajectory rather than a snapshot, which means a wobbly day 30 followed by a solid day 60 reads as normal learning instead of a red flag.
Why open-ended evidence beats quiz scores
A quiz tells you someone can recognise the correct policy. A written response tells you whether they can apply it to a messy case, and it tells you how they think while doing it. Those are different capabilities, and the second is the one that predicts on-the-job performance.
Open-ended evidence also fails usefully. A wrong multiple-choice answer tells you only that it was wrong. A written paragraph shows whether the person misunderstood the policy, applied it correctly to the wrong facts, or understood it and expressed it badly. Each needs a different intervention.
The objection is cost: judging written work at scale is slow, which is the real reason most programmes settle for quizzes - a resourcing constraint dressed up as a design choice.
Rubrics for onboarding competencies
Open-ended evidence is only defensible if you judge it against stated criteria. Three rules do most of the work.
Anchor each criterion to something observable in the work. "Shows good judgement" is unmarkable. "Identifies the escalation trigger in the scenario and states who it goes to" is markable, and two managers will agree on whether it was met.
Keep the criteria count low. Four to six per checkpoint. Onboarding rubrics bloat fast because everyone wants their concern represented, and a fifteen-criterion rubric gets skimmed rather than applied.
Define what "not yet" looks like. Most onboarding rubrics only describe success, which leaves the failing end of the scale to interpretation - and the failing end is where the consequences sit. We go through criteria count, level descriptors, weighting and the classic failure modes in rubric design for workplace competency assessment.
Where AI evaluation fits
AI is genuinely good at the bottleneck: reading a large volume of written or spoken responses and producing a first-pass judgement against a rubric with the supporting evidence quoted from the response. That is the task that makes open-ended onboarding assessment affordable at cohort scale rather than for ten people.
It is also good at consistency and at surfacing patterns. The fortieth response gets the same rubric applied the same way as the first, which is not true of a tired reviewer, and if eleven of forty new starters missed the same criterion, that is a curriculum problem rather than eleven individual problems.
Where it does not
AI should not make the call on anything with consequences attached. A day 90 result affecting probation is a decision about someone's employment, and a person needs to own it - both because solely automated decisions of that kind carry weight under GDPR and because a manager who cannot explain the reasoning cannot defend it.
It also cannot see context. The new starter who wrote a thin answer because they spent the week covering an absence looks identical on paper to the one who did not engage. Their manager knows the difference; the model does not.
And it should not be trusted where your rubric is vague. A model given a fuzzy criterion will produce a confident score anyway, which is worse than a human hesitating. If the rubric would not survive two managers disagreeing about it, fix the rubric before adding automation. On the broader question of how far to trust the scores, see is AI grading reliable.
Closing the loop
The output should be two things: a specific development conversation for each person, and a change to the programme. If nothing about your onboarding changes after a cohort has been through it, the evaluation was theatre. Cohort-level reporting is what makes the second half possible - not "average score 72" but "criterion four was missed by a third of the cohort, and here are the responses", which points straight at the week of onboarding that needs rewriting.
Where Scorafy fits
Scorafy scores open-ended responses - written, audio, video, uploaded files - against your rubric, quoting the evidence for each criterion, and routes every result to a named reviewer who can override any score before release. Cohort reports show where a group collectively fell short, not just individual results, and nothing reaches a respondent until a human releases it.
It is not an HRIS and does not run your onboarding programme. It handles the evaluation layer - the part where someone has to read forty responses and judge them consistently. If that is the bottleneck stopping you measuring capability instead of satisfaction, see how scoring works or the AI assessment software overview.