Start with what is actually being measured
Before any question about AI, ask the older question: what does this assessment measure, and what is the evidence that it measures it? An AI can score a badly designed assessment with perfect consistency and the results will still be meaningless. Look for tools that make your assessment design explicit - your criteria, your performance levels, your weighting - rather than tools that score against criteria you cannot see. If the vendor claims the assessment predicts job performance or competence, ask for the validity evidence behind that claim, not a description of the AI that scores it.
How the AI scores - the six properties that matter
- Rubric alignment: the AI scores against YOUR criteria and performance levels, in your language, not a hidden generic notion of quality. If you cannot read the rubric the AI used, you cannot defend the score.
- Evidence citations: every criterion score comes with the specific part of the answer that earned it. A score without evidence is an opinion with a decimal point.
- Consistency: the same answer should get the same score. Ask how the vendor constrains variability between runs and across a cohort, and how you would detect it if consistency slipped.
- Human review and override: a reviewer sees the AI draft before the candidate does, can change any individual criterion score, and the override is recorded with who, when, and why - not just an approve/reject button on the whole report.
- An audit trail: the rubric as it stood at scoring time, the AI’s scores and evidence, every human change, and who released the result - retrievable months later without archaeology.
- Drift monitoring: AI models change. Ask what happens when the vendor updates the model behind your assessments, whether you are told, and how anyone would notice if scoring behaviour shifted mid-cohort.
Two of these deserve their own reading: what genuine human review looks like versus a rubber stamp, and what a complete audit trail contains.
Bias and fairness: evidence, not claims
Every vendor will tell you their AI is fair. The useful question is what they can show. Ask whether they have tested scoring outcomes across demographic groups, on what data, and what they found - including the uncomfortable parts. Ask how the design limits the AI's exposure to irrelevant signals: an AI that scores a written answer against a rubric has less room to absorb bias than one scoring video, voice, or personality signals, but less room is not no room. And ask what the process is when a candidate alleges a biased result: if the answer does not involve a human re-scoring from the recorded evidence, the fairness story has no enforcement mechanism.
The generative-AI era: your candidates have AI too
Any written take-home assessment now has to assume some candidates will draft answers with ChatGPT or similar. Be sceptical of vendors selling AI-text detection as the answer: detection tools produce enough false positives that an accusation built on one is difficult to defend, and enough false negatives that a clean result proves little. Better questions for a vendor: what assessment formats do you support that are harder to outsource - timed tasks, structured follow-ups, video or audio responses, file-based evidence of real work? Can the assessment disclose to candidates upfront what tools are and are not permitted? A tool that helps you design around the problem is worth more than one that promises to catch it after the fact.
Security and privacy: the verifiable answers
- Residency: where is candidate data stored and processed, can you choose the region, and does the AI inference itself stay in-region or leave it?
- Retention and deletion: how long are responses kept, is deletion on request real and complete, and what happens to media files like video answers?
- Training use: are your candidates’ responses used to train the vendor’s or anyone else’s AI models? The answer should be a flat no, in the contract, not a settings page.
- Sub-processors: the full list, including the AI model provider - because that is where the responses actually go - with notice before it changes.
- Attestations: what independent evidence exists (SOC 2, ISO 27001, penetration tests)? Where the answer is “not yet”, what can they show instead - and do they say so unprompted?
Candidate experience
The people being assessed did not choose the tool, and their experience is part of what you are buying. Check whether candidates are told an AI is involved before they start - in plain language, not a privacy-policy footnote. Check what the assessment feels like to complete: does it work on a phone, does it save progress, what happens on a dropped connection during a timed section? And check what candidates receive at the end. A tool that returns specific, evidence-based feedback turns the assessment into something of value to the person who sat it; a tool that returns a bare number leaves them guessing.
Integration and administration
The unglamorous criteria that determine whether the tool survives contact with your organisation: how respondents get in (shareable links, email invitations, CSV import of a whole intake), how results get out (CSV and PDF export, webhooks, an API), team roles and permissions, and what your identity requirements are - if your organisation mandates SSO, put it on the checklist early, because not every vendor in this category has it. Same for LMS integration if results need to land in Moodle, Canvas or similar: ask specifically, and ask what "integration" means - a genuine LTI connection and a CSV upload are very different answers wearing the same word.
The vendor evidence pack to demand
Before a paid commitment, ask for a pack of artefacts rather than a slide deck: documentation of how scoring works, end to end; any accuracy or human-agreement figure with its method published alongside it; a sample audit record for a single completed assessment, showing the rubric, the AI's evidence, an override, and the release; the security documentation and sub-processor list; and a reference or case study you can verify. Where a vendor has no agreement figure yet, the honest alternative is a supported trial: measure the agreement yourself on your own rubric and your own respondents. That number, on your material, is worth more than any vendor's benchmark anyway.
Total cost per completed assessment
Price lists in this category are hard to compare because the unit varies: per seat, per assessment, per response, per credit. Normalise everything to one number - total cost per completed, reviewed, released assessment at your real volume - and include the costs that do not appear on the pricing page: reviewer time per result (transparency that speeds review is a cost lever, not a luxury), setup and rubric-building time, and any per-feature surcharges for things like video responses or extra team members. Then check what happens at the edges of your plan: what a burst month costs, and whether unused capacity rolls anywhere.
Questions to ask any vendor - including us
These are the questions we think a buyer should put to every tool in this category. We answer them for Scorafy below, including the ones where our answer today is not yet.
- Can I see the exact rubric the AI scored against, and the evidence behind each score? - Scorafy: yes. You author the rubric, and every criterion score carries the quoted evidence that placed it.
- Can a reviewer change one criterion score, and where is that recorded? - Scorafy: yes. Overrides are per-criterion with a recorded comment, and the AI’s original score stays on the record beside the reviewer’s.
- What stops a result reaching a candidate unreviewed? - Scorafy: nothing reaches a respondent until a person releases it; the release is stamped. For certification-style decisions the system prompts for confirmation and records it.
- What is your published accuracy or human-agreement number? - Scorafy: not yet. We publish the measurement method we recommend, and we support running the comparison on your own cohort - but we do not yet have an independent published figure, and we will not invent one.
- Do you have SSO? - Scorafy: not yet. Authentication is email and password today. If SSO is a hard requirement, we are not currently the tool for you, and we would rather say so here than in month two.
- Do you integrate with my LMS? - Scorafy: not yet. There is no LTI integration today; results move via CSV and PDF export, webhooks, and a read-only API.
- Is candidate data used to train AI models? - Scorafy: no. Responses are processed to generate the evaluation and are not used to train models.
- Where does the data live? - Scorafy: EU hosting as standard, with an Australian environment for organisations that require it. The sub-processor list is published on the security page.
Common questions
- What should I look for in AI assessment software?
- Six things, roughly in order: evidence the assessment measures what it claims to measure; transparency about how the AI scores, including the evidence behind each score; genuine human review with score-level override and a recorded reason; a complete audit trail; bias and fairness evidence rather than assurances; and security and privacy answers you can verify - data residency, retention, whether your data trains models, and the sub-processor list.
- How can I tell whether an AI scoring tool is accurate?
- Ask the vendor for a published accuracy or human-agreement figure together with the method behind it: what was compared, on whose rubrics, over how many scripts, and how disagreement was counted. A number without a method is marketing. If no figure exists, run your own check - have the AI and your own assessors score the same set of real responses independently and measure the agreement before you rely on it.
- Should AI assessment software make the final decision on a result?
- No. The defensible pattern is AI drafts, a person decides: the AI proposes criterion-level scores with cited evidence, a reviewer with authority confirms or overrides each one, and the result reaches the candidate only when a person releases it. A tool that sends AI-scored results straight to candidates has removed the step that makes the outcome yours to defend.
- What security and privacy questions should I ask an AI assessment vendor?
- Where the data is stored and processed (region, and whether you can choose); how long responses are retained and whether deletion is real; whether your candidates’ responses are used to train AI models; the full sub-processor list including the AI provider; and what independent attestations exist - and if the answer is none yet, what the vendor can show instead.
- How should AI assessment tools handle candidates who use generative AI themselves?
- Honestly. AI-text detectors are unreliable enough that accusations built on them are hard to defend, so be wary of any vendor selling detection as a solved problem. Stronger designs reduce the incentive instead: assessment formats that require personal or observed evidence, timed and structured tasks, oral or video follow-ups, and clear policies candidates agree to before starting.