Controlled, reviewable AI assessment
The AI that marks the answers.
It reads each response against your rubric, cites the lines it scored on, and holds every result until your assessor releases it.
Start free1 assessment · 10 evaluations a month · no card
Try it here
Watch it read one real answer.
This is a written response to a delegation question. Press score, and the sentences it marks on light up.
Response, question 3
I plan the week on a Monday and set out who is covering which account. If something slips I would rather find out early than late, so we do a short stand up on Wednesday.
I tend to keep the difficult handovers myself rather than pass them on, because I know the client and it feels faster than briefing someone properly. When the team is under pressure I will take work back off them to protect the deadline.
Scored against your rubric
An illustration. Nothing is sent anywhere.
The honest starting point
Two assessors, one answer, two different marks.
What already happens
On a borderline result, judgement fills the gap the rubric left.
This is the ordinary state of marking written work. Two capable assessors read the same answer and land in different places, the difference is rarely written down, and nobody can reconstruct it a year later.
What gets asked for later
The completed work, and the reasoning behind the mark.
Reviews and audits sample real responses rather than policies. So the question is not whether a machine can be trusted to mark. It is whether you can show your working when somebody asks to see it.
Scorafy does not remove the assessor. It hands them a first draft with the evidence already attached, and it writes down what happened next. How scoring works.
The pipeline
How a score gets made.
Every mark on Scorafy travels the same eight stages. You can inspect all of them.
Answer
A respondent writes, speaks, records, or uploads. 26 question types, including video and audio with transcription, and PDF or Word documents.
Rubric
Your criteria and levels - up to 10 criteria across 6 levels, with per-question weighting and your own grading schema.
AI evaluation
Claude by Anthropic reads the actual response against your rubric. Not a score-to-template lookup.
Evidence
Every score cites the specific answers behind it, so you can see why - not just what.
Confidence
The model reports how certain it is about each question, so your reviewer can go straight to the marks it considers borderline.
Human review
Your assessor reviews each result, overrides any score, and adds comments. Overrides are recorded.
Release
Respondents see nothing until the assessor releases the result. The release, and who made it, is on the record.
Report
A unique, evidence-backed report for each respondent - and one cohort report across the whole group.
Assessment should be fast and fair. Scorafy does the reading and the first draft. Your assessor keeps the judgement.
See the pipeline run on a real answer.
Start free1 assessment and 10 AI evaluations a month, free. No card.
Live sittings
Run the whole room at your pace, not theirs.
Not everything is a link you send out and wait on. When you need a cohort assessed together - in a classroom, a workshop, an exam hall - Scorafy runs an assessor-paced sitting where you control what the room can see, and the server backs you up.
Available on every plan, including free.
You open the questions
One at a time, when the room is ready. Everyone answers the same question at the same moment - no one races ahead, no one reads question nine while you are still briefing question one.
Enforced on the server
An unopened question is not sent to a candidate’s browser at all, and the platform refuses an answer to a question you have not opened. It is a boundary, not a hidden div.
Brief, pause, discuss
Drop instruction screens in to set up a section, and pause-and-discuss beats between questions when you want the room talking rather than typing.
Join by link or QR
Put the QR code on the screen and the room joins in seconds. No accounts, no installs, no roster wrangling on the day.
Nothing is lost
Answers save to the server incrementally as candidates type. A flat battery, a dropped connection or a closed lid does not cost anyone their work.
Rehearse it first
Run the whole sitting against yourself before the room arrives, so the first time you drive it is not in front of thirty people.
Prefer respondents to work in their own time? That still works exactly as it always has - self-paced links, timers, and scheduling are unchanged. Live sittings are an extra gear, not a replacement.
Built for accountability
Designed for organisations that need controlled, reviewable AI assessment.
Not promises - controls. Each one is live in the product today, and described in full on our security page.
Release gate
Results are held until a human assessor reviews and releases them. Respondents see nothing - not a score, not a report - until sign-off.
Server-enforced timers
Set a limit for the whole assessment, per question, or both. Enforcement happens on the server, not in the browser - an expired timer is rejected even if the page is tampered with. The clock starts only when the respondent presses Start.
Overrides on the record
Assessors can override any AI score, per question, with a comment. The original AI score is kept alongside the human decision.
Audit trail
Who scored what, when it was overridden, and when it was released. The trail is the point.
One-click audit pack
Export an assessment’s full evidence chain as PDF and JSON - model and configuration versions, the override ledger, and the release chain - in a single file a moderator or auditor can read without a login.
Confidence on every mark
Each question comes back with the model’s own certainty alongside the score. Reviewers sort by it and review the borderline marks first, instead of reading every result at the same depth.
Segregation of duties
The person who builds an assessment need not be the person who releases its results. Roles are separable, and every release records who made it.
Retention you control
Set response retention per assessment, from 30 to 365 days. Personal data is removed on your schedule, not ours.
Data residency
EU hosting (Dublin) as standard, with an Australian residency option (Sydney) for organisations that need data onshore.
MFA available. Row-level security on every table. Evaluations run on Claude by Anthropic, and your data is never used to train AI models. Full detail: /security · How scoring works
Proof, honestly
What we can show you, and what we won't claim.
On the record
- Every score arrives with its evidence. Open the example report and check.
- The full scoring pipeline is documented, in plain language, at How scoring works.
- Our security posture is described as controls, not badges, at /security.
- A public API (v1) with webhooks - including a
report.releasedevent that fires only after human sign-off - is documented at /docs/api. - We won't publish an industry accuracy figure - but we do measure agreement inside your own organisation. Your Accuracy page shows how often your assessors agree with the AI, which questions attract the most overrides, and how each assessor moves scores. Your number, from your own reviewed data.
- If AI capacity runs out mid-cohort, a submitted response is parked and evaluated when capacity returns. A respondent who has submitted is never failed by our meter.
What you won't find here
- No headline accuracy percentage. We won't publish one until we can back it with a proper human-agreement benchmark - the method we'd have to meet is public at measuring AI-assessor agreement.
- No certification badges we haven't earned.
- No stock-photo testimonials. When customers say something we're allowed to quote, we quote them by role, verbatim.
If a claim on this page isn't linked to something you can inspect, tell us and we'll remove it.
Pricing
Pricing that scales with your cohorts
Start free. Scale to cohort and enterprise volume when you're ready.
Free plan available
1 assessment, 10 AI evaluations per month. No card needed.
Starter
For solo assessors and small teams
- 3 active assessments
- 100 AI evaluations per month
- All 26 question types
- Conditional branching
- CSV export
Growth
For teams assessing multiple cohorts
- Everything in Starter, plus
- 15 active assessments
- 500 AI evaluations per month
- Custom branding
- PDF report export
- Analytics dashboard
- 10 team members
Pro
For firms and training providers
- Everything in Growth, plus
- Unlimited assessments
- 2,000 AI evaluations per month
- Unbranded assessments and reports
- Weighted rubrics with evidence
- API access & webhooks
- Unlimited team members
- Priority support
For organisations
Business & Enterprise
From $1,000 /month, billed annually
Multi-instructor accounts, data residency options, configurable retention, and dedicated onboarding for teams running assessments at scale.
Talk to usAll prices in USD. One AI evaluation is one completed respondent assessment, evaluated. Verified schools and non-profits get 40% off Starter, Growth, and Pro.
Questions
Everything you might be wondering
“No tool let us add our own criteria and rubrics for assessment. This is exactly what we needed.
Instructional Designer, Global Education NFP
One AI evaluation is one completed respondent assessment, evaluated. It is counted when someone completes your assessment and Scorafy generates their report. Partial or abandoned attempts do not count towards your limit.
No. Respondents click a link and start the assessment. No signup, no friction. If an assessment is timed, respondents see a start screen first - the clock only starts when they press Start - and if they return after finishing they see a completion screen rather than dropping back into the questions.
When your assessor releases them. Scorafy holds every result - score, evidence, and report - until a human reviewer signs off. You can override any score before release, and overrides are kept on the record.
Yes. Set a time limit for the whole assessment, individual questions, or both. Limits are enforced on the server, so they hold even if a respondent's browser is manipulated.
We use Claude by Anthropic. Every report is generated fresh from the respondent's specific answers - not matched to pre-written text buckets. The AI's score is a first draft - your assessor reviews and releases every result. See how scoring works at /how-scoring-works.
Yes. In an assessor-paced live sitting you open one question at a time and everyone in the room answers it together, with instruction screens for briefing and pause-and-discuss beats in between. Sequencing is enforced on the server - the content of a question is not sent to anyone’s browser until you open it, and the platform refuses answers to questions that have not been opened yet. Candidates join by link or QR code, and live sittings are available on every plan, including free.
Every mark carries the model’s own confidence for that question, and every criterion carries an explanation of which rubric level was awarded and why. Reviewers work the low-confidence marks first rather than reading every result at the same depth. Your organisation also gets an Accuracy page showing how often your assessors agree with the AI, which questions attract the most overrides, and how each assessor moves scores.
Yes, if you want that. An assessment can be set to share the written feedback immediately while holding the numeric score back until an assessor signs it off. The score is withheld server-side - it does not reach the browser, and it is kept out of the AI’s written prose too.
Yes. You can add context, frameworks, rubrics, and grading schemas. The AI uses these to generate domain-specific reports that sound like your practice.
No. All plans allow unlimited questions. We recommend 8 - 25 for the best balance of respondent experience and AI analysis quality.
All data is encrypted in transit and at rest, with row-level security isolating every organisation’s data, and MFA available on all accounts. You control response retention per assessment, from 30 to 365 days. We never use your data to train AI models. Full detail on our security page at /security.
See it live
See what your reports could look like.
Answer five questions in the interactive demo and read the evidence-backed report it generates - then imagine it reviewed, signed off, and released by your team. Takes about a minute.
Free plan included · No credit card required