Workplace competency rubrics fail in predictable ways. They have too many criteria, the level descriptors are adjectives, the weighting was set by feel, and there is no clear definition of what a fail looks like. Each of those is fixable, and fixing them changes marker agreement more than any amount of assessor training.
This is about structure rather than wording - how many criteria, what a level descriptor needs to contain, whether to weight, and what to do about the failure modes. If you want the writing-level guidance, our piece on how to write an assessment rubric covers phrasing criteria so they can be applied consistently.
How many criteria
Four to seven for most workplace competencies. Below four you are usually bundling distinct capabilities into one criterion. Above seven, two things go wrong at once: markers stop applying the later criteria properly, and the criteria start overlapping, so a single flaw in the work gets penalised three times.
That double-counting effect is the one people miss. Take a rubric with criteria for "communication", "clarity of written explanation", and "audience awareness". A candidate who writes densely fails all three for the same underlying reason, and their overall score now says they are weak across a broad range when the truth is one specific weakness. The rubric has amplified a single fault into a pattern.
The test for whether two criteria are genuinely separate: can you picture a submission that clearly passes one and clearly fails the other? If not, merge them.
If a competency genuinely needs twelve things assessed, that is a signal for two assessment tasks rather than one twelve-criterion rubric.
How many levels
Three or four. Workplace competency often only needs two - competent or not yet competent - and there is nothing wrong with that where the decision itself is binary. Where you want developmental information, three levels work well: not yet, competent, and a level above that describes genuinely stronger practice.
Five and six-level scales look more precise and are usually less reliable, because nobody can articulate the difference between level three and level four in a way two markers apply identically. If you cannot write a sentence distinguishing adjacent levels in terms of the work itself, that boundary does not exist.
Level descriptors: describe the work, not the worker
This is the single highest-leverage change in most rubrics. A descriptor should say what the submitted work contains, not how good the person is.
Weak, and extremely common:
- Excellent - demonstrates a thorough understanding of risk management.
- Good - demonstrates a sound understanding.
- Developing - demonstrates a limited understanding.
That scale is one adjective repeated. Thorough, sound and limited are not observable. Two markers will place the same submission differently and both will be able to justify it, which means the rubric is not doing any work.
The same criterion, rewritten around the work:
- Competent with strength - identifies all four risks present in the scenario, states a mitigation for each that addresses the specific risk named, and identifies which risk to act on first with a reason.
- Competent - identifies all four risks and states a workable mitigation for each. Prioritisation may be absent or unexplained.
- Not yet competent - misses one or more of the four risks, or proposes a mitigation that does not address the risk identified.
Now the boundary between levels is a countable fact about the submission. Markers agree because there is nothing left to interpret. This is also what makes the criterion gradeable by a machine with evidence cited against it - a model can point to where the four risks were addressed, or where the fourth was not.
A second worked example, for a customer-facing role:
- Competent - the response states what went wrong in plain language, says what will happen next with a timeframe, and does not use internal terminology without explaining it.
- Not yet competent - omits the next step or the timeframe, or uses internal terminology unexplained, or attributes blame to the customer.
Notice that the "not yet" descriptor lists the specific ways to fail rather than saying "does not meet the standard". That is deliberate, and it is the next section.
Define the fail band explicitly
Most workplace rubrics describe success in detail and leave failure as the absence of it. That is backwards, because the fail band is where the consequences and the appeals live.
Three things a fail descriptor should settle:
- What is missing versus what is wrong. A submission that omits a required element and one that includes it incorrectly are different failures. Say whether they land in the same place.
- How absence is handled. If the candidate never addressed the criterion at all, is that the lowest level or is it unassessable and returned? Systems that guess at this produce inconsistent results.
- Whether any single criterion is fatal. In safety-critical and regulated work, some criteria are non-negotiable - a strong submission that misses the mandatory safety step is not a pass regardless of the total. If that is true for you, it must be written into the rubric as a rule, not left to marker judgement.
Weighting: useful sometimes, a smokescreen often
Weighting lets you say criterion one matters more than criterion five. That is legitimate when it is genuinely true and you can say why.
Two cautions.
First, weighting can compensate in ways you did not intend. If criterion three is a mandatory safety behaviour and it is weighted at 15%, a candidate can fail it and still pass overall on the strength of the other criteria. If a criterion is essential, it needs a minimum standard rule, not a heavy weight. Weight expresses importance to the total; it cannot express "this one is required". Those are different mechanisms and conflating them is a real source of indefensible outcomes.
Second, weighting is often used to paper over a criteria-count problem. When a rubric has eleven criteria and six of them are weighted at 3%, those six are not being assessed in any meaningful sense - they are decoration that costs marker time. Delete them.
A reasonable default: equal weights unless you can state the reason for the difference in one sentence. If you cannot, the difference is not real.
The four failure modes, in order of frequency
- Vague descriptors. Adjectives instead of observable content. Fix by rewriting each level in terms of what appears in the work. This one change usually does more for marker agreement than everything else combined.
- Too many criteria. Leads to skimming and to double-counting a single fault. Fix by merging overlapping criteria and splitting the task if it genuinely needs more.
- No fail band definition. The rubric describes success and leaves failure implicit. Fix by writing the fail descriptor as specifically as the pass descriptor, and by settling how absence is treated.
- Bundled criteria. "Clear and well-evidenced" forces one score onto two independent qualities. A clear, unevidenced answer has nowhere to sit. Fix by splitting - one idea per criterion.
Test the rubric before you deploy it
Take six previously-marked submissions, deliberately weighted toward borderline cases, and have two assessors apply the new rubric independently. You are not checking whether they reach the right answer. You are checking whether they reach the same answer, and where they diverge.
Every disagreement points at a specific criterion or level boundary. Rewrite that one and run it again. A rubric that produces agreement on your hardest cases will hold up in front of an auditor, an appeal panel, or an AI grader - all three are testing the same property, which is whether the criteria mean one thing. It is also the precondition for automating anything: a vague rubric produces confident nonsense at scale. The rubric is the work; the tooling is downstream of it.
Where Scorafy fits
Scorafy applies your rubric to open-ended responses and quotes the evidence from the submission behind each criterion score, then routes it to a reviewer who can override anything before release. Rubric versions are recorded against results, so a rubric you improve next quarter does not retroactively change what past learners were judged against.
The rubric stays yours. The tool applies it consistently and shows its working - it does not decide what competent means. If you want to see a rubric applied with cited evidence on your own material, see how scoring works or the product guides.