AI in Learning and Development

Does a bigger AI model score people better? We tested six.

Field Note 02 ยท 26 August 2026

First, what this test was, in one minute

Skip this if you have read can AI assess soft skills?

A certification run is a conversation, scored not by asking an AI for a mark out of 100 but by asking it small, specific questions about specific moments. Did they name the issue. Did they stay calm when pushed. Yes, partly, or no. Arithmetic adds those up.

Any AI model can be put in the scorer's seat. So the question this entry asks is a fair one: does it matter which one? To find out, you need a human to compare against. One person reads the same conversations and answers the same questions, and you measure how often each model lands where the human landed.

Think of football referees. Give six referees the same rulebook and the same 36 matches, and compare their calls against a senior referee's. You are not asking which referee is famous. You are asking which one calls the game the way the senior official does, and whether the rulebook is doing its job. That is the whole of this entry.

The assumption we started with

Before a score is allowed to decide whether someone earns a credential, we have to know it means something. So we did the ordinary thing, and it produced a result we did not expect.

We took 36 recorded practice conversations in one skill, difficult feedback. One person rated all 36 blind: shuffled, unlabelled, no way of knowing which had been designed to be strong and which to fail. That rater was our founder, and it is a limit we state everywhere these numbers appear. Then we ran six AI models over the same transcripts, from three different vendors and across the whole range of price and size, and compared each one against the human.

Here is the sensible assumption, and it is the one almost everybody holds. A bigger, more expensive model reads more carefully, so it should agree with an expert human more closely. Buy the best referee you can afford. That is the assumption most of us start with, whether we are buying an assessment tool or building one in-house: pick the best model you can afford, and the scores will follow.

And at first glance the results seemed to bear out at least the first half of it. Every one of the six cleared the bar the testing profession publishes for trusting an automated scorer in high stakes exams. Every one of them separated the strong conversations from the weak ones perfectly. Six competent referees.

Then we looked at the order, and the assumption fell over. The two cheapest models agreed with the human best. The flagship agreed with the human least, and it marked nearly twice as harshly as the best calibrated model in the group. The most famous referee blew the whistle most.

And no single model won everything. The models that tracked the human's ranking most closely were not the ones that landed closest to his actual scores: the one nearest his numbers, less than two points off, came last of the six on agreement. There was no overall winner to crown, which turned out to be the useful part.

The human consequence is the one to hold onto. If you are building an AI assessment, in-house or for sale, the name of the model inside it tells you nothing about whether its scores agree with a person. It tells you what it cost. Only a measurement against a human tells you the rest.

Here is what did not fail. Every model, cheap or expensive, told a strong conversation from a weak one every single time. The method held. What varied was only how closely each model tracked one human's judgement, and how far below that human it landed.

Then we ran it again, and the order reversed

We had scheduled a re-run to check for drift, on the principle that a published figure with no expiry date is not a figure. Nine weeks later we scored the same 36 conversations, with the same checks, against the same human ratings, at the same settings. Nothing on our side changed.

The model that came last of the six on agreement now came first. One of the two cheapest models, which had been joint first, now came last. The flagship was no longer the weakest fit; it had been overtaken from below. And the two models that used to sit closest to the human's scores had crossed over and now sat slightly above him rather than below.

Why, in one sentence: the referees had been retrained between the two seasons, and nobody had told us.

AI models are updated by the companies that make them, sometimes under the same name. Of the six we had recorded in June, one identifier no longer answered at all and was being served by a differently named build, and two we had noted as retired were still answering. A vendor's own lifecycle notices are not a reliable account of what your code is actually calling, which is why we now record which identifier answered on every single call.

None of that is a story about one vendor beating another. It is a story about what kind of fact a model ranking is. Ours had a shelf life of nine weeks, and it expired without anyone telling us. If we had not re-run it we would still be publishing the June order, and it would still look exactly as authoritative as it did on the day we measured it.

The correction that expired with it

There is a practical edge to this, and it cost us something.

Here is the sensible idea. If you know a referee is consistently a little harsh, you can measure by how much and allow for it. The model that scores our certifications marked about four points harsher than the human in June, so it carried a fixed correction to compensate. Measure the harshness once, add it back, and the scores line up. On the June figures that is exactly what it did.

On the re-run the same model marked slightly above the human. The correction had not become slightly wrong. It had become wrong in the direction that hands learners points they did not earn. A calibration constant is a measurement with an expiry date, and ours had quietly expired. We are rebuilding it from the retained output.

What did not fail is worth saying. The re-run caught it. That is what a scheduled re-measurement is for, and it is the reason we had one.

What we did not do is correct the harshness across the board with a blanket per-model offset, which was the obvious fix and which our own experiment argued against.

And the one thing that did not move at all: all six models still sit within a hair of each other on agreement, across three vendors and a wide range of price and size. That is the claim this entry was actually about, and it is the claim that survived.

The wrong lesson

"Cheap models win" is the shareable version. We do not believe it, because the same programme refuted it three times: twice from the other side at the time, and once more when the ranking itself reversed nine weeks later. It was the more quotable half of this entry and we are not going to leave it standing on that basis.

One: the cheap tier lost the job it was actually doing. Separately from scoring, we tested which model should write the checks a course assesses against. Same brief, three arms, one model each. The newest mid tier model produced a set where 7 of 15 checks in the final assessment could be passed by reciting the rubric back instead of doing the skill, and two rounds of automated repair did not fix it. The top tier model wrote the set that was hardest to game, and it was also the fastest. We moved authoring up to the top tier, and the single largest change in that decision was taking rubric writing away from one of the cheap models.

Two: the cheap tier failed in a place we had not thought to look. Our fastest, cheapest tier could not convincingly play a well trained learner when we used it to generate practice transcripts for validation. As a scorer, the same tier gave full marks to an answer that recited the rubric rather than demonstrating anything. A model that cannot be fooled by a stuffed answer is doing a different job from a model that agrees with a human about a genuine one, and being cheap does not tell you which of those two jobs a model is good at.

What we actually think is going on

The method carries the validity. Not the model.

A score here is not one model's opinion of a conversation. The conversation is broken into small, specific checks at engineered moments. Each check is a yes, a partly, or a no. Arithmetic does the weighing. That structure is why a cheap model can hold its own against an expensive one: we never ask any model for an overall verdict. We ask it a series of narrow questions it can actually answer, and we keep the weighing out of its hands. The rulebook does the work. The referee only has to read it.

Which turns model choice from a status contest into a casting decision. Every model earns one specific seat: the one that scores, the one that writes the checks, the one that plays the learner when we test our own instrument. The model that lost the scoring seat here is the one we now pin as the validation instrument, because stable and slightly severe is exactly what that seat wants.

We do not swap the scorer casually, because changing it means recalibrating everything ever measured against it. Certification today is scored by a single model, and the jury of six is a capability we built and measured rather than a safeguard currently running. It would be easy to buy the biggest model available and put its name on the front page. We would rather publish the seating chart and the numbers behind it.

What this does not prove

Everything above rests on one rater. One person read all 36 conversations and scored them, and every agreement figure is agreement with that one person. What we have not measured is how much two trained people agree with each other on the same conversations. Until we do, we cannot say how much of the gap between a model and a human is the model. Multi rater work is running now and we will publish it whichever way it lands.

It is also one skill and 36 conversations. Difficult feedback, nothing else. We make no claim that these figures carry to another skill before we have run another skill. And the conversations were written to a design rather than recorded from live learners, which is a real limit on what the numbers describe: they tell you how the scorers behave on material built to be scored.

The authoring result is a single run per arm. It was enough to move a model, it is not a measurement, and we said so in our own record before acting on it. None of this is a ranking of AI models in general. It is how six models behaved on one task, in one configuration, on one set of days.

The referee's rulebook

Go back to the six referees. Nobody sensible picks a referee by fee and then never checks their calls. You check them against a senior official, on real matches, and you check again next season, because the referees you hired have been on courses since and you were not in the room. What made the calls consistent was never the referee. It was a rulebook specific enough that any competent official reading it lands in the same place.

That is what we built, and it is why the answer to "which model?" is a smaller question than it looks. The rulebook is ours. The referees are hired, they change, and we re-check them on a schedule. A ranking of them is true on the day it was measured, and we say which day.