Why assessment teams specifically should care
Most AI Act commentary is written for AI vendors. But the Act regulates deployment as well as development, and its high-risk list names the two contexts assessment lives in: education and vocational training (including AI used to evaluate learning outcomes) and employment (including AI used to evaluate people for recruitment, promotion and performance). If AI contributes to scores that affect a person's qualification, certification, hiring or progression, the Act is about you, not just your vendor. Its obligations have been phasing in since 2025, with the high-risk regime continuing to phase in through 2026 and 2027 - which makes now the right time to set the operating pattern, not after an audit letter.
The Article 4 duty: competent humans, not just compliant systems
Article 4 is short and easy to underestimate: organisations must ensure the people operating AI systems have a sufficient level of AI literacy. For an assessment operation that translates concretely. Assessors need to understand what the AI actually evaluated, read the evidence it cited, know its failure modes - fluent-sounding answers that dodge the criterion, thin evidence dressed as confidence - and know when and how to override. An assessor who rubber-stamps AI scores they do not understand is an Article 4 gap wearing a workflow.
Transparency: the respondent finds out first, not last
The Act's transparency provisions - and plain fairness - point the same way: people should know AI is involved in assessing them before they are assessed. The workable pattern is a disclosure on the assessment's opening screen: AI will analyse the answers, a human reviewer can override any AI-suggested score, and results are not final until released. Burying it in terms nobody reads, or disclosing after results land, fails the purpose even where it might scrape past the letter.
Human oversight that would survive an audit
The Act expects oversight of high-risk AI to be real: a person with the competence, information and authority to intervene. In assessment terms that means the human sees the AI's draft before the respondent does, can change any score with a recorded reason, and controls when a result is released. For certification and other consequential decisions, the system should prompt for confirmation and record it - so the record shows a named person decided, with the AI's contribution and the human's decision both documented. Oversight that leaves no record is indistinguishable from no oversight, which is why the audit trail is the compliance backbone rather than a nice-to-have.
The deployer's checklist
- Classify deliberately: write down what your AI-assisted scores influence (grades, certification, hiring, progression) and whether that puts the use in the high-risk categories - do not leave scope to assumption.
- Disclose before assessment starts: AI involvement, the human review right, and that results are released by a person.
- Keep the audit records: model versions, configuration snapshots, evidence citations, override ledger, release records, exclusions.
- Evidence Article 4 literacy: train assessors on reading AI evaluations and overriding them, and let the override ledger show real engagement.
- Measure agreement: track how often assessors change AI scores and by how much - oversight without measurement is a ritual.
- Check the data layer separately: GDPR runs in parallel (lawful basis, retention, data residency), and the AI Act does not absorb it.
- Put appeals in writing: a respondent who disputes a score should trigger reconstruction from the record, not archaeology.
None of this requires slowing assessment down. It requires the process to leave evidence - which is what turns "we use AI responsibly" from a claim on a slide into something you can show.
Common questions
- Is AI-assisted assessment high-risk under the EU AI Act?
- It can be. Annex III captures AI used to evaluate learning outcomes in education and vocational training, and AI used to evaluate people in employment contexts such as recruitment, promotion and performance. Whether a specific deployment lands in scope depends on what the output influences - which is a question to settle deliberately, not by default.
- What is Article 4 of the EU AI Act?
- Article 4 requires providers and deployers of AI systems to ensure a sufficient level of AI literacy in the staff who operate them - people using AI need to understand what it does, where it fails, and how to exercise oversight. For assessment teams, that means assessors who understand what the AI scored, what evidence it cited, and when to override it.
- Do respondents have to be told AI is involved in assessing them?
- Transparency obligations in the Act point that way for systems interacting with people, and it is good practice regardless: tell respondents before they start that AI analyses their answers and that a human can review and override the scores. Disclosure after the fact is not disclosure.
- Can AI make certification decisions under the Act?
- The Act’s design pushes consequential decisions toward meaningful human oversight. The defensible pattern for certification is that the AI drafts an evidence-cited evaluation and the platform prompts a person for confirmation and records it - so the decision is a human’s, evidenced by the record.
- Does the AI Act apply to organisations outside the EU?
- It reaches any organisation whose AI system’s output is used in the EU, not only EU-registered companies. An Australian RTO or a US company assessing EU-based staff or learners should assume relevance and take the checklist seriously.