Before a score is allowed to decide whether someone earns a belt, we have to know the score means something. So we did the ordinary thing, and the ordinary thing produced a result we did not expect.
We took 36 recorded practice conversations in one skill, difficult feedback. One person rated all 36 blind: shuffled, unlabelled, no way of knowing which conversation had been designed to be strong and which had been designed to fail. That rater was our founder, and it is a limit we state in every place these numbers appear.
Then we ran six different AI models over the same 36 transcripts, scoring them the same way, and compared each model against the human.
We expected the ranking to follow the price list. Bigger model, closer agreement. That is the assumption the market runs on. It is the assumption sitting behind every "powered by" badge on a competitor's page.
What came back
The two cheapest models in the set agreed with the human best.
The flagship model was the weakest fit on every agreement measure we ran, and it marked nearly twice as harshly as the best calibrated model in the group.
And no single model won everything. The models that tracked the human's ranking most closely were not the models that landed closest to the human's actual score. The model that sat nearest the human's numbers, less than two points off, came last of the six on agreement. There was no overall winner to crown, which turned out to be the useful part.
One thing every model did equally well: all six separated the strong conversations from the weak ones perfectly. Not one of them confused a good conversation for a bad one. The spread was never about whether a model could tell the difference. It was about how closely it tracked a human's judgement, and how far below the human it landed.
All figures below: n = 36 transcripts, one skill, one human rater, corpus written by AI, six scorer models, ratings collected blind.
| Scorer | Weighted kappa | ICC | Score bias vs human |
|---|---|---|---|
| cheapest tier, model A | 0.87 | 0.95 | 6.2 points below |
| cheapest tier, model B | 0.87 | 0.96 | 5.4 points below |
| flagship model | 0.84 | 0.93 | 9.2 points below |
| closest calibrated model | 0.81 | 0.95 | 1.9 points below |
Weighted kappa measures how often two raters land on the same judgement, corrected for the agreement you would get by chance, and it counts a near miss as less wrong than a wide miss. ICC is the intraclass correlation coefficient: how much of the variation in scores is real differences between learners rather than disagreement between raters. Both run 0 to 1. Published thresholds for qualifying an automated scorer in high stakes testing sit at 0.70 for both. Bias is the average signed gap in score points.
Full jury of six against the human: kappa 0.87, ICC 0.96. That ICC interval is not exactly recoverable, because the raw scoring outputs from the run were lost. An approximation from the point estimate alone gives 0.92 to 0.98, and we label it approximate wherever we quote it rather than dropping the interval and publishing a bare number.
Strong versus weak discrimination: AUC 1.000 on all six models.
Every bias in the set is negative. Models mark the same learner below the human, consistently. That is a scale offset, not a disagreement about who did well.
The wrong lesson
"Cheap models win" is the shareable version. We do not believe it, because the same programme refuted it twice from the other side.
One: the cheap tier lost the job it was actually doing. Separately from scoring, we tested which model should write the checks a course assesses against. Same brief, three arms, one model each. The newest mid tier model produced a set where 7 of 15 checks in the final assessment could be passed by reciting the rubric back instead of doing the skill, and two rounds of automated repair did not fix it. The top tier model wrote the set that was hardest to game and it was also the fastest. We moved authoring up to the top tier. The single largest change in that decision was taking rubric writing away from one of the cheap models, which had been doing it until then.
Two: the cheap tier failed in a place we had not thought to look. Our fastest, cheapest tier could not convincingly play a well trained learner when we used it to generate practice transcripts for validation. As a scorer, the same tier gave full marks to an answer that recited the rubric rather than demonstrating anything. A model that cannot be fooled by a stuffed answer is doing a different job from a model that agrees with a human about a genuine one, and being cheap does not tell you which of those two jobs a model is good at.
One run per arm. Exploit and repair counts carry real run to run variance and we recorded that caveat in the source before we ruled on it. We treated it as evidence strong enough to move a model, and not as a measurement. The scoring jury was deliberately left untouched by that decision.
What we actually think is going on
The method carries the validity. Not the model.
A score here is not one model's opinion of a conversation. The conversation is broken into small, specific checks at engineered moments. Each check is a yes, a partly, or a no. Arithmetic does the weighing. That structure is why a cheap model can hold its own against an expensive one: we never ask any model for an overall verdict. We ask it a series of narrow questions it can actually answer, and we keep the weighing out of its hands.
Which turns model choice from a status contest into a casting decision. Every model has to earn one specific seat: the one that scores, the one that writes the checks, the one that plays the learner when we test our own instrument. The model that lost the scoring seat in this test is the one we now pin as the validation instrument, because stable and slightly severe is exactly what that seat wants. And we do not swap the scoring jury casually. Changing the scorer means recalibrating everything that was ever measured against it.
There is a version of this business that buys the biggest model available and puts its name on the pricing page. We would rather publish the seating chart and the numbers behind it.
What this does not prove
Everything above rests on one rater. One person read all 36 conversations and scored them, and every agreement figure is agreement with that one person. What we have not measured is how much two trained people agree with each other on the same conversations. Until we do, we cannot say how much of the gap between a model and a human is the model. Multi rater work is running now and we will publish it whichever way it lands.
It is also one skill and 36 conversations. Difficult feedback, nothing else. We make no claim that these figures carry to another skill before we have run another skill. And the conversations themselves were written to a design rather than recorded from live learners, which is a real limit on what the numbers describe: they tell you how the scorers behave on material built to be scored.
The authoring result is a single run per arm. It was enough to move a model. It is not a measurement, and we said so in our own record before we acted on it.
None of this is a ranking of AI models in general. It is how six models behaved on one task, in one configuration, on one set of days. Models change underneath you. Ours already have, which is the next thing.
The gate we hold ourselves to is the published criteria used to qualify automated scorers in high stakes testing: weighted kappa and correlation at 0.70 or better against expert human rating, with a bounded standardised mean difference and a bounded degradation from the human benchmark, measured per scorer model, re measured on a schedule, published, and corrected by calibration when a model drifts out of band. Those are not our thresholds. They are the field's.
The unmeasured quantity named above is inter rater reliability. Everything here is agreement against a single rater, so no degradation figure in this entry is benchmarked against a human to human baseline. That is the gap the multi rater study exists to close.
Since then
Dated 26 August 2026.
The authoring models moved to the top tier and stayed there. The scoring jury has not been changed, on purpose, because changing it means a recalibration pass.
We are now scheduling a re run of the same six models on the same 36 transcripts to see whether the numbers have moved. We already know part of the answer: some of the model versions used in the original run have been retired by their vendors since, so "the same models" may not resolve any more. That is not a footnote about the study. It is the finding. A score you published a quarter ago is only as stable as the model underneath it, and nobody tells you when that model changes. Watching for that drift is becoming a product feature rather than an experiment we repeat.
We also decided not to correct the harshness with a blanket per model offset, after our own experiment argued against it. That decision has its own entry.
For where the numbers stand today, see the validity study (forthcoming). For how the scoring method got built at all, see the score that wandered. For what happened when the instrument caught its own defect, see the narration trap.
Ask any vendor for their human rater agreement numbers. These are ours, with the n and the limits attached.
