First, how our marking works, in one minute
You need three things to follow the rest of this, and they are all simple.
A certification run is a conversation. The learner talks to a character who pushes back on them, the way a real customer or a real direct report would. Nothing is multiple choice.
The conversation is not given one overall grade. It is broken into small, specific questions about observable behaviour. Did they name a date. Did they acknowledge the concern before answering. Did they hold their position without escalating. We call each of those a check, and a course carries a few dozen of them.
Each check gets one of three answers: yes, partly, or no. Yes is worth full marks, partly is worth half, no is worth nothing. Add those up, turn it into a score out of a hundred, and compare it to a pass mark. Ours is 75.
That is the whole machine. Everything below is about one number that was supposed to sit between the marking and the pass mark, and did not.
The promise we made to ourselves
Here is the problem the missing number was meant to solve, and it is a real one.
Different markers are differently strict. Put the same conversation in front of two experienced assessors and one will give it 70 and the other 78, not because either is wrong but because people carry different bars in their heads. The AI models that do our marking are no different from each other in this respect. One of ours marks noticeably harder than the other.
If you know one marker runs harsh, the obvious remedy is to compensate. Work out by how much, then add that much back to everything they mark. Our own build plan said exactly this. It specified a global adjustment for each AI grading model, and it called that adjustment a floor, because the point of it was that it always exists. No course anywhere would ever be marked without at least that much correction applied.
We wrote that. We believed it. It sounds like the responsible thing to do, and for two months nobody in the building questioned it.
The afternoon we went to check
In July we went to look at what the floor was actually doing to real certificates. This was housekeeping, not suspicion.
The answer was that it was doing nothing at all.
The adjustment existed. It was being calculated, correctly, every time. But the code that awards certificates had been rebuilt some months earlier, and the rebuilt version reads the marking straight from the AI grader and then looks up a different table for its corrections. That second table is meant to hold corrections specific to each course, and it had never had anything put in it. It was empty.
So the sequence went: mark the conversation, look up the correction, find nothing, apply nothing, compare the uncorrected score to the pass mark, award or refuse the certificate.
Every certificate we had ever issued had been decided on raw, uncorrected marking. The safety net described in our own documents was hanging in a room nobody walked through.
Two people held certificates at that point. Both had passed comfortably, far above the line, so neither result was affected and nobody was harmed. But that was luck, not design. If either of them had landed near the pass mark, the decision would have gone one way when our own written policy said it should have gone the other, and we would not have known.
What everybody does when a marker is too harsh
So we went to attach the safety net properly. And the fix we were about to build is the one almost everybody reaches for, which is why it is worth walking through slowly.
If one teacher in a department marks harder than the others, the standard remedy is to add a few marks to every paper they touch. Work out that they run about four marks below their colleagues, add four marks back, done. You have not changed how they mark. You have simply compensated for a known bias, and everyone they marked is now on the same footing as everyone else.
That is what our floor was. Four marks, added to everything one AI model touched, seven for the other. It was specific, it had measured values behind it, and the work would have taken an afternoon.
Then we sat down and worked out what adding four marks to every paper actually delivers, and the answer is not four marks. It is not close to four marks, and the amount it is short by depends on who the candidate is.
Why padding does the opposite of what you expect
Here is the part that surprised us, and it comes down to one very simple fact.
A mark that is already full cannot take a bonus. There is nowhere for it to go.
Our correction was applied check by check, not to the final score. So take a candidate who demonstrated a behaviour completely. That check is already at full marks. Add four to it and it is still at full marks, because full marks is the top of the scale. The bonus is simply discarded.
Now take a candidate who half demonstrated that behaviour. That check sits at half marks, there is plenty of room above it, and the full bonus lands.
Follow that through to its conclusion. A strong candidate, who scored full marks on most of their checks, has almost nowhere to put the correction, so they receive almost none of it. A weak candidate, who half landed nearly everything, has room everywhere, so they receive nearly all of it.
We had built a correction that is largest exactly where the performance is weakest.
That has two consequences and both are bad.
The first is that it squashes the scale. The gap between your strongest and weakest performers narrows, because the weak one is being lifted and the strong one is not. Separating those two people is the entire reason you bought an assessment. We would have been paying for a correction that undoes the thing we were selling.
The second is worse, and it is about where decisions get made. Nobody cares what a correction does to someone scoring 95 or someone scoring 20. Those outcomes are not in doubt. The only place the number matters is right at the pass mark, where one or two marks decide whether a person walks away certified. And a candidate sitting at the pass mark is, by definition, somebody with a mixture of full marks and half marks. So they receive some of the correction and not all of it.
Our four mark correction delivered under two marks to the only candidates whose result it could change.
The two candidates who swapped places
That was enough to stop us. What came next is what made the decision easy.
Padding does not only squash the scale. It can reorder it. Two candidates can change places, and the one who performed worse can end up passing while the one who performed better fails.
Picture two of your people.
The first is a coaster. They half landed almost everything. They touched every topic, made a gesture at every behaviour, and committed fully to almost none. Their marking is a long run of half marks with very few outright misses and very few full marks.
The second is genuinely stronger. They demonstrated most behaviours completely and properly. They also missed a handful outright, because they went deep rather than wide and ran out of conversation.
On raw marking the second person scores higher. Anyone reading the two transcripts would agree with that, and so would the marking.
Now pad both papers.
The coaster's paper is almost entirely half marks, which means it is almost entirely room. Every single check takes the full bonus. The stronger candidate's paper is mostly full marks, which means it is mostly ceiling, and most of their bonus evaporates the moment it is applied.
The coaster gains far more from the correction than the stronger candidate does, purely because they left more space to be topped up. Add enough padding and the coaster passes and the stronger candidate fails.
Nothing about either performance changed. Nothing about the marking changed. The ranking inverted because of an arithmetic property of a correction we had designed to be fair.
Say that out loud in front of a works council. Explain that the person who genuinely performed better was refused a certificate, and the person who skated through got one, and that the reason is a compensation factor in the scoring engine. That is not a defect you argue about in a design review. That is one you lose.
Then our own data said the harshness was not even where we thought
At this point we had a mechanism that misbehaved. What we did not yet know is that the thing it was correcting for had been misdiagnosed from the start.
The whole premise of a per-model correction is that harshness belongs to the model. This AI model marks hard, that one marks soft, so correct each one by its own amount. It is a statement about the markers.
A few weeks earlier, for entirely unrelated reasons, we had run an experiment on the questions instead.
Some of our checks had been written badly. They asked two things in one breath. Did the learner acknowledge the concern and offer a next step. A question like that has no honest single answer when a learner does one and not the other, so a marker has to invent a rule for themselves, and different markers invent different rules. We took one of those double questions and split it into two single questions.
That one change removed almost all of one AI model's harshness by itself. Not reduced it. Removed it. The model that had appeared to be marking two points too hard turned out to be marking correctly, and to have been tripped up by a badly written question.
So the harshness was not a property of the marker. It was a property of how we had written the check.
And that matters enormously here, because checks are written per course. If one course happens to contain three badly written checks and another course was written carefully, a correction derived from the first course is meaningless in the second. Apply it globally and you are taking a wording problem that belongs to one course and awarding marks for it in courses that never had the problem. You are not correcting anything. You are handing out free marks, evenly, to people who did not need them.
What the measurement world already knew
We were not the first people to discover this, which was both reassuring and a little embarrassing.
There is a body of research on exactly this question, and its finding is blunter than ours. A marker is not uniformly strict. The same marker is harsher on some questions and more lenient on others, at the same time. Strictness is a pattern, not a single setting.
Sit with that for a second, because it kills the whole idea of a single correction number. If a marker runs two marks harsh on some checks and two marks soft on others, then a single correction of plus two is wrong in both directions simultaneously for the same person. It over-corrects half their marking and under-corrects the other half. It does not matter how carefully you measure it, because there is no single right value to measure. You are looking for a number that does not exist.
The formal name for this, if you ever want to look it up, is differential rater functioning. We had been building a remedy that the field had already shown cannot work.
Why we did not just go and get a better number
There is one more salvage attempt available at this point, and we considered it seriously. Keep the correction, but measure it much more carefully and derive a better value.
We did not, and the reason is the most uncomfortable thing on this page.
Every figure had come from a single marker. One person, marking generated practice transcripts, in one skill. When you compare an AI's marking against a human's and call the difference bias, you are treating the human as the truth. With one human, you have no way of knowing whether that human is representative, strict, lenient, or simply having an off week.
This is not a small caveat. With one marker there is no such thing as a margin of error, because there is nothing to compare them against. You cannot say the true value is four plus or minus one. You can only say that one person's marking produced four. The field that does this professionally, setting pass marks for exams that decide careers, typically wants around fifteen judges before it trusts a standard.
So a more carefully measured correction would have been a more precise number sitting on top of an anchor we had no way to check. It would have looked earned, and it would not have been, which is a worse position than an obviously provisional one, because a number that looks earned invites you to trust it.
And that limit has not gone away since. No customer's panel has set a bar on their own people's real transcripts yet. It is the largest open exposure in this system, it is not closed by anything described on this page, and it closes when the first panel sits. Two of our own colleagues have marked a live course independently, which is how we know the panel mechanism works with real people at a distance, but four shared checks between two markers is nowhere near a validity result and we will not present it as one.
So we withdrew the promise, and here is what that cost
We did not build the floor. We deleted the promise from the plan instead of keeping it badly.
That decision has a price, and the price falls on real people, so it needs saying plainly rather than buried.
Our pass mark of 75 was set at a time when everybody assumed a correction was coming. It was chosen to sit sensibly on corrected marks. Marks are now compared to it uncorrected, which means the bar is harder to clear than the number was designed to be. Our certification runs strict, deliberately, and we know by roughly how much.
Being strict means the mistakes we make run in one direction. We refuse certificates to people who had actually earned them. We do not hand certificates to people who had not.
That is the direction to be wrong in, and we would choose it again. A certificate awarded to someone who cannot do the job is worthless to everyone who holds one, and it is the failure that destroys a credential permanently. A certificate withheld from someone who deserved it is a serious thing, but it is recoverable: they sit it again.
But it is a choice, not an accident, and the people it costs are real. So it sits in our architecture record with the reasoning attached, and our learner-facing wording is not permitted to describe the bar as calibrated, because it is not.
The rule that came out of it, and you can use it anywhere
Here is the thing worth taking away even if you never look at our product.
A correction that is the same for every question is not really a correction to the scores at all. It is a change to the pass mark wearing a disguise.
Watch the arithmetic. Add four marks to everybody's paper and pass people at 75. Now instead change nothing and pass people at 71. Those two things produce precisely the same list of passes and failures. Every single candidate is in the same category either way. It is the same decision, reached by two routes.
Except one route has a scoring adjustment nobody can see, which has to be documented, defended and maintained, and which can reorder your candidates on the way. And the other route is a single honest sentence: our pass mark is 71.
So do the honest version. Move the cut, not the scores.
There is an important exception, and it is the whole of our current approach. A correction that differs from question to question is a completely different animal. If check one gets adjusted by two marks and check four by nothing and check seven by five, you have genuinely reshaped the marking, and no change to the pass mark can reproduce that. That kind of correction is real and it belongs on the scores.
But it can only be built from real human judgement about each specific check, which is why our correction table sits empty rather than full of guesses. It is not empty because we forgot. It is empty because nothing has earned a place in it yet.
The part where we were protected by an accident
Six weeks later we went to look at how that table gets filled, and found that our careful decision had been protected by nothing but luck.
Here is how filling it is supposed to work, and the design is sound. An expert from the customer's own business reviews the machine's marking, check by check. Where they disagree with it, that disagreement is recorded. Once enough disagreements accumulate on a particular check, the system works out how far the machine is drifting on that specific check and starts correcting it. Human judgement goes in, a correction comes out, and it is specific to the check rather than global.
That is exactly the kind of correction the rule above permits, and we had described it as switching itself on gradually as evidence built up. Which sounds prudent.
"Enough disagreements" turns out to be five. Five reviews on one check and the system starts adjusting live certification decisions.
And there is a second thing. The material those reviewers were being shown was not real learner work at all. It was our own generated practice samples, built deliberately to span the range from excellent to deliberately terrible, because their job was to test whether a check can tell good from bad. A set of samples built to span a range is not distributed anything like a real cohort of your employees. A correction fitted on one tells you very little about the other.
Put those together. Five reviews of artificial samples on a single check would have begun moving real certification decisions, quietly, using exactly the kind of unearned number we had just spent two months refusing to invent.
The only thing standing between us and that was the fact that nobody had opened the review queue yet.
So the lesson here is not really about calibration at all. A decision that is protected only by nobody using a feature is not protected. It is a coincidence with a deadline on it. We moved that behaviour behind a named constant in the code rather than a switch in a settings screen, so that turning it on requires reading the reasoning for why it is off.
What to ask anybody selling you assessment
Three questions, ours included. They cost nothing and the answers tell you a great deal.
Ask whether their pass mark is a number or a judgement. A pass mark that emerged from a model, or is a round number, or was settled on because the results looked about right, is not a standard. Ask who decided it, what they were picturing when they decided, and whether you can read that reasoning anywhere. If the answer is that it has always been that, you are looking at a habit rather than a standard.
Ask what happens when their grader marks too hard, and listen for the word adjustment. If scores get corrected by a fixed amount, ask two follow-ups. What does that adjustment do for a candidate who is already scoring near full marks. And can two candidates change places because of it. The answer to the second should be no, and they should be able to show you why rather than assure you.
Ask what is switched off today, and what is holding it off. Every assessment system has something in it that has not yet earned the right to be live. A good answer names the thing and names the mechanism keeping it disabled. The worrying answer is that it is fine because nobody has turned it on yet, which is not a control. It is a coincidence.
We got that third one wrong in our own system and did not notice for six weeks. It is on this page because a research log containing only the parts we got right is an advertisement.
For how the scoring instrument was built in the first place, see the score that wandered. For the experiment that showed a badly written check defeats any marker, see the obvious fix was half wrong. For the time our own checklist rewarded describing the skill over doing it, see the narration trap.
One marker, one skill, generated transcripts. The bar is provisional and we say so. SkillDojo. Practice makes proof.
