A course had gone live with certification switched on. Then the results started arriving and they did not make sense.
One learner scored 58 against the 75 needed to pass. She had covered every part of the conversation, and her certification answers were stronger than her practice answers, which had scored as high as 88.
A second tester went in deliberately, playing the roleplay with full knowledge of exactly what the grader was looking for. Quoting the buyer's own words back. Making explicit conditional trades. Naming documents. Best result available to him: 72. Still a fail.
And one area of the assessment scored 50 while the written feedback under it told the learner their answer had been correct.
The obvious reading is that the bar was set too high. That was our first theory too. It was wrong, and finding out why took one evening and one database query.
How we found it
We pulled the 13 individual checks from the failing 72 run out of the database and read each one against its verdict and against the grader's own written explanation of that verdict.
The pattern was not subtle.
The 8 checks he passed all asked the same kind of question: did the learner do the thing. Did they attach a condition to the concession. Did they acknowledge the concern before answering. Did they hold their position without escalating. Concrete, observable actions. He did them. He got full marks.
The 5 checks that cost him the belt were different in kind, not in degree. Four of them required the learner to explain their own reasoning out loud, and the grader's own explanations gave the game away. On three of those four, the grader agreed that the learner had exercised the skill, and then withheld the mark because he had not said so explicitly. The phrasing in the record is "you implied it, but did not explicitly state it." Another said the learner had responded without any sign of analysing what he was responding to.
No competent salesperson stops mid negotiation to lecture the buyer on how to tell a genuine buying signal from a polite one. Someone reciting a rubric passes checks like that. Someone actually doing the job fails them.
The fifth was worse. It asked the learner to point at a specific signal of polite but non committal interest. The scenario had cast the buyer as genuinely interested. That signal was never in the conversation. He was scored zero on a moment that never happened.
Our own written model says, in as many words, that nothing should be scored that the conversation never gave the learner a chance to show. The rule was right there. The defect walked in through a side door we had not thought to watch.
And that dissolved the mystery of the 50 with the affirming feedback. An area's score is the average of its checks. One full mark, two half marks, and one impossible zero average to exactly 50. The feedback the learner reads quotes the passing check's encouraging explanation, while the half marks and the zero drag the number down silently. It was not the grader being inconsistent. The checklist itself was mis measuring.
Why every defence we had waved it through
This is the part worth your time, because the failure was not that we lacked a quality gate. We had one, and it passed this course.
Every course goes through a publish gate that tests whether each check can actually tell a strong answer from a weak one. It works by generating sample answers at both levels and confirming the check separates them.
The strong sample answers are written by a model. And that model is shown the checks while it writes them.
Read that twice, because we did not, for months. A sample author who can see a check asking whether the learner articulates a criterion will naturally write an answer that articulates the criterion. So a narration check separates strong from weak perfectly on those synthetic answers, sails through the gate, and then fails every real human who talks like a human.
Our gate proved that a rubric aware strong answer passes. It never once checked that a naturally speaking strong performer passes. The blind spot was not in the checks. It was in the shape of the test we used to validate the checks.
The class of error is a validation corpus that shares information with the instrument it is validating. The separation statistic we relied on, area under the curve, was genuinely high and genuinely meaningless here, because both arms of the comparison were generated with sight of the marking scheme. AUC measures separation. It cannot measure whether the thing being separated resembles a real learner. Any evaluation harness where the reference answers are produced from the rubric carries this exposure, and it will not show up in the metric that harness reports.
What we changed
Three things, in the order we did them.
Repair the live instrument. Eight checks across two areas of that course were rewritten from narration to demonstration: from asking whether the learner named a distinction, to asking whether they acted on it, with something a listener could actually observe. The unanswerable one was rebuilt into a question the learner can be asked in any version of the scenario.
Stop the generator minting more of them. Two rules were added to the instructions that write new checks. First: a check must be passable by doing the behaviour in natural speech, and must never require the learner to state, name, or show awareness of their own reasoning, unless explaining is itself the skill being taught. Second: only ask about something the scene is guaranteed to stage. If the roleplay cannot reliably produce the moment, the check has no business existing.
Give the gate an honest control. The publish gate now also generates answers as a strong performer who has never seen the marking scheme, facing the same moments a live learner faces. A check that rubric aware answers pass and natural answers fail gets flagged as rewarding narration rather than skill.
The proof
Same day. A natural certification run, with no rubric narration anywhere in it, scored 90 against the 75 cut. The belt was issued. The certificate rendered in its designed form for the first time this course had ever produced one.
And while we were walking the fixed version, the new control caught a third defective area live. It had been scoring learners 38 for not explaining why a phrase they had correctly responded to signalled readiness. Same defect class, an area nobody had flagged, found by the thing we had just built to find it.
Two smaller fixes rode along, both of which should have existed already. One grader returning malformed output can no longer take down a whole belt; the jury scores over its surviving members. And the grader is now explicitly forbidden the contradiction we had seen with our own eyes: if your written explanation says the behaviour was fully demonstrated, the verdict is yes.
We left the pass mark exactly where it was. A genuinely strong run now clears 75 with 15 points of margin, which is the evidence that says the instrument was the problem and the bar was not.
The two things that made this harder than it sounds
Neither of these was in the plan, and both changed the answer.
The blind control had to be a graduate, not an expert. The first live runs of the new gate flagged the checks about conditional concessions. But making a trade instead of giving ground away free is not rubric trivia. It is the actual lesson of that course. The control was failing it because a generically skilled seller who never took the course does give ground away free.
Three people at the same moment make it concrete. One has read the marking scheme and performs to it. One is a skilled seller who never took the course. One is a skilled seller who did take the course, where the discipline is taught. Real certification candidates are the third person. A control built as the second person would have pushed authors to weaken every taught discipline just to get a course published, and eroded the bar to "sounds like a good salesperson."
So the control was rebuilt to role play a graduate of the course: it gets the syllabus, the skill briefs, the concept names, everything a real learner legitimately carries out of the teaching, and never the marking scheme. Syllabus in, marking scheme out.
Then we had to demote our own new gate. One session later the flag went from blocking publication to advising the author. Fixing the control's identity had not fixed its actor. The model playing the graduate could not reliably embody a taught discipline in natural conversation, and it whipsawed across repeated runs on the same course.
The reasoning for demoting it is the same reasoning that built it. If a synthetic tester is not reliably strong enough to demonstrate a taught skill, then letting it block publication means rewording sound checks to satisfy it. That is exactly the bar erosion we had just spent a session avoiding. So the flag now tells the author what it found and never acts on its own. Discrimination and resistance to gaming stayed as blocking bars, because those are deterministic and this is not. Promoting it back is a one line change the day a strong enough model makes the signal trustworthy.
What the gate does now, in both directions
We already had a defence against a learner who is worse than they sound: the checks are built to resist someone stuffing an answer with rubric language they do not understand.
What we did not have, until this, was a defence against a learner who is better than they sound. Somebody who does the job well and does not narrate it.
A certification that only catches one of those is not measuring skill. It is measuring fluency in our own vocabulary.
What this does not prove
This was found on two learners on one course, at the point of discovery. That is not a sample, and we are not claiming a rate. What it establishes is that the defect class exists, that it survives a quality gate designed to catch bad checks, and that our published validation numbers could not have seen it, because those numbers were measured on material written with the marking scheme in hand.
The fix is the same shape as the finding: we changed how checks are written, and we changed the test that validates them. Whether the class is fully gone is something only live use tells us, and we found the third instance of it by walking a course we had just declared fixed.
The certification bar itself has a separate story. Our graders do run harsher than a human rater, and we made a deliberate decision not to correct that with a blanket offset. That decision, and the experiment that argued against the obvious fix, is a forthcoming entry.
Since then
Dated 26 August 2026.
The two authoring rules ride the generation instructions verbatim and the repaired checks stand. The naturalness control is still advisory, still surfaced to the author on every publish, and still waiting on a corpus model strong enough to be trusted with a veto.
We have not seen a fourth instance of the class since the third was caught, which we read as encouraging and not as evidence.
For how the scoring instrument was built in the first place, see the score that wandered. For why the most capable model was the worst scorer, see method beats model size. For where the validity numbers stand today, see the validity study (forthcoming).
A bar only means something if you can fail it. It also has to be failable for the right reasons.
