How scoring works
From answer to signed-off report
Scorafy doesn’t create assessments - it evaluates human answers. Here is the pipeline every response goes through, where the humans stay in control, and how we measure whether the machine is earning its place in it.
- 01
A person answers your questions
Open-ended answers in their own words - long-form, short-form, or transcribed from audio and video. This is the raw material: what the person actually said, not which option they clicked.
- 02
Your rubric frames the evaluation
Your criteria, your performance levels, your weightings - per assessment or per question. The AI is never asked "is this good?"; it is asked "where does this answer sit against this rubric?". Question weights carry through to the final arithmetic.
- 03
The AI evaluates the actual answer
Claude reads each answer against the rubric and your methodology context. Answers are isolated from instructions - content inside an answer cannot steer the evaluation. Every evaluation records the model used and what it consumed.
- 04
Every judgement cites its evidence
Each rubric criterion comes back with the level selected and the evidence for it, quoted from the respondent's own answer. A score you cannot trace to the answer is not a score you can defend - so every score is traceable.
- 05
The model says how sure it is
Alongside each question’s score, the model reports its own confidence in that judgement, and each criterion carries an explanation of which level was awarded and why. This is the part most tools leave out: a reviewer should not have to read thirty results at the same depth to find the three that need them. Sort by confidence and the borderline marks come to the top.
- 06
A human reviews and signs off
Assessors see the draft evaluation first. They can override any score with a comment, and nothing reaches the respondent until an assessor releases it. The release is stamped and audited - the AI drafts, a person decides.
- 07
The released report goes out
Only after release does the respondent receive their result - the report, the feedback, and the grade if you use grading schemas. An assessment can also be set to share the written feedback straight away while holding the numeric score back until an assessor confirms it. Cohort-level reporting then aggregates across respondents for the programme view.
- 08
And then you measure the pipeline itself
Every override your assessors make is a data point about whether this pipeline is working. Your Accuracy page turns them into a number: how often your assessors agree with the AI, which questions attract the most overrides, and how each assessor moves scores. We will not publish an industry accuracy figure we cannot defend - but you get yours, measured on your own reviewed data.
Why this holds up under scrutiny
Rubric in, evidence out, a named human sign-off in between, and an audit trail underneath. That chain - not “the AI is accurate” - is what makes an AI-assisted result defensible to a moderator, an auditor, or the person being assessed.