First, how the test works, in one minute
Skip this if you have read our other field notes.
The test is a conversation. The learner talks to an AI character that plays the other side of a real work situation. In this course, that was a buyer in a sales negotiation. The whole conversation is recorded, and then a second AI marks it by answering a list of small, specific questions about what happened. Did the learner acknowledge the buyer's concern before answering? When they gave something up, did they ask for something in return? Each question is marked yes (full marks), partly (half marks) or no (nothing).
The questions are grouped into sections. A section's score is the average of its questions, out of 100, and the learner sees it with a short written comment from the AI marker explaining the score. The pass mark is 75.
Think of a spoken language exam. A fair examiner marks whether you can hold the conversation: order the meal, handle the complaint, ask for directions. They do not mark whether you can explain the grammar you just used. The moment an examiner starts wanting the explanation, fluent speakers begin to fail and people who memorised the textbook begin to pass. That is what happened to us. It took one evening to find, and we had missed it for months.
The results that did not make sense
We launched a course with certification switched on. When the results came in, they did not add up.
One learner scored 58, well short of the 75 needed to pass. She had covered every part of the conversation, and her answers in the real test were stronger than in her practice rounds, which had scored as high as 88.
A second tester, a colleague of ours, took the test knowing exactly what the AI marker was looking for. He quoted the buyer's own words back to them. Every time he gave ground, he asked for something in return. He named the documents he would send. The best he could get was 72. Still a fail.
And in one section the score was 50, while the AI's written comment underneath told the learner their answer had been right.
The obvious explanation was that the pass mark was too high. That was our first thought too. It was wrong, and finding the real cause took one evening and one look at the data.
How we found it
We took the 72 apart question by question. For each of the 13 questions we read the mark the AI gave and the reason it wrote down for that mark.
The pattern was obvious.
The 8 questions he passed all asked the same kind of thing: did the learner do it? When they gave something up, did they ask for something back? Did they acknowledge the concern before answering? Did they hold their position without raising the temperature? These are actions you can hear in the conversation. He did them, and he got full marks.
The 5 questions that cost him the certificate were a different kind of question altogether. Four of them required the learner to explain their own thinking out loud, and the AI's written reasons gave it away. On three of those four, the AI agreed the learner had used the skill, and then held back the mark because he had not said so. In its own words: "you implied it, but did not explicitly state it." Another marked him down for responding to the buyer without showing any sign of analysing what the buyer meant.
This is the language examiner wanting the grammar. No good salesperson stops in the middle of a negotiation to explain to the buyer how they can tell real interest from polite interest. Someone reciting the mark scheme passes questions like that. Someone actually doing the job fails them.
The fifth was worse. It asked the learner to spot a sign that the buyer was being polite but not committing. But the scenario had written the buyer as genuinely interested, so that sign never appeared. He scored zero for a moment that never happened, like a language student marked down for not using a tense the conversation never gave them a reason to use.
Our own design rules say, in plain words, that nothing should be marked if the conversation never gave the learner a chance to show it. The rule was there. The problem came in through a door we were not watching.
That also explained the 50 with the positive comment. A section's score is the average of its questions. One full mark, two half marks and one impossible zero average to exactly 50. The comment the learner saw quoted the encouraging reason from the one question they passed, while the half marks and the zero pulled the number down without a word. The AI was not being inconsistent. The questions themselves were measuring the wrong thing.
Why our quality check missed it
This is the part worth your time, because we did have a quality check, and this course passed it.
Here is the sensible idea behind it. Before a course can go live, every question has to prove it can tell a strong answer from a weak one. So our system writes sample answers, some strong and some weak, and confirms that each question passes the strong ones and fails the weak ones. A question that cannot tell them apart is sent back. That is a reasonable thing to build, and it had caught bad questions before.
It failed on one detail. The strong sample answers were written by an AI that could see the questions while it wrote them.
Read that twice, because we missed it for months. If the writer can see a question asking whether the learner explains their reasoning, it will write an answer that explains its reasoning. So a question that rewards explaining passes every strong sample, sails through the quality check, and then fails every real person who talks like a real person.
It is an exam board testing its spoken exam against model answers written by the person holding the mark scheme. Of course the questions looked fine. The one thing those model answers could never show was whether a fluent speaker who had never seen the mark scheme would pass, because no such answer was ever tested.
What that meant for people: for weeks, the learners this certification failed were the ones doing the job well. Our quality check proved that a strong answer written with the mark scheme in view passes. It never checked that a strong performer speaking naturally passes.
What did not fail matters too. The 8 questions that asked whether the learner did something you could hear were sound, and the AI marked them correctly. The AI marked consistently, and its written reasons were honest enough to expose the problem. And the pass mark was right where it should be. The blind spot was not in the marking or the standard. It was in how we tested the questions before they went live.
What we changed
Three things, in this order.
We repaired the course. Eight questions across two sections were rewritten from asking whether the learner explained something to asking whether they did it, in a way a listener could actually hear. The question about the sign that never appeared was rebuilt into one the learner can answer in any version of the scenario.
We stopped the AI that writes our questions from making more like them. Two rules were added to its instructions. First, a question must be passable by doing the skill in natural speech. It must never require the learner to state, name or show awareness of their own reasoning, unless explaining is itself the skill being taught. Second, only ask about something the scene is guaranteed to include. If the roleplay cannot reliably produce the moment, the question should not exist.
We gave the quality check an honest tester. It now also writes answers the way a real graduate of the course would, and we had to think carefully about who that is.
Picture two experienced salespeople. Both are good at their jobs. The first has never taken the course. Years of selling have given them habits that mostly work, and one of them is common: when a deal stalls, they give a little ground, a small discount or an extra, to keep it moving, and ask for nothing back. That habit is exactly what this course teaches people to break. Its core lesson is never to give something up without getting something in return. So the first salesperson fails that question, and they should. The test is working.
The second salesperson is just as experienced, but took the course, practised the new habit using roleplays and made it part of how they sell. That is who a real learner is, and that is who the tester now plays. It knows everything the course teaches. It has never seen the mark scheme.
The difference matters. If the tester played the first salesperson, it would flag the course's most important lessons as faults, and fixing those "faults" would mean watering down everything we teach.
When a question passes the answers written with the mark scheme but fails the graduate's natural answers, the system warns the course author that the question may be rewarding explanation instead of skill. The author then decides whether to rewrite it. The warning does not stop the course from going live on its own, and that is deliberate. The AI playing the graduate is not yet consistent: run it twice on the same course and it can show the skill well one time and poorly the next. If a tester that inconsistent could stop a course from going live, we would end up rewriting good questions just to please it.
The proof
The same day, a strong performer took the test again, speaking naturally and reciting nothing from the mark scheme. The score was 90, against a pass mark of 75. The certificate was issued, the first this course had ever produced.
And while we were checking the fixed version, the new tester caught a third faulty section, live. It had been scoring learners 38 for not explaining why a phrase they had responded to correctly showed the buyer was ready to move forward. The same kind of fault, in a section nobody had flagged, found by the tool we had just built to find it.
Two smaller fixes came along, both of which should have existed already. If one of the several AI models that mark a test returns broken output, it can no longer sink the whole certificate; the test is scored by the models that did answer. And the AI marker is now explicitly forbidden the contradiction we saw with our own eyes: if its own written reason says the learner fully did it, the mark is yes.
We left the pass mark exactly where it was. A genuinely strong performance now clears 75 by 15 points, which tells us the test was the problem, not the standard.
Protecting both kinds of learner
We already had protection against a learner who sounds better than they are: the questions are built to resist someone stuffing an answer with words from the mark scheme they do not understand.
What we did not have, until this, was protection for a learner who is better than they sound. Somebody who does the job well and does not talk about how.
A certification that protects against only one of those is not measuring skill. It is measuring how well people speak our vocabulary.
What this does not prove
We found this with two learners on one course, at the moment we discovered it. That is not a sample, and we are not claiming how often it happens. What it does show is that this kind of fault exists, that it can pass a quality check designed to catch bad questions, and that our published validation numbers (how closely our AI agrees with a human marker) could not have seen it, because those numbers were measured on answers written with the mark scheme in hand.
Nor can we promise this kind of fault is gone. Only real use will tell us that. We found the third case in a course we had just declared fixed.
The examiner's sheet
Go back to the spoken language exam. Nobody would defend an examiner who fails a fluent speaker for not explaining the grammar, and nobody would defend an exam board that checked its questions only against answers written by the person holding the mark scheme. Both are easy to see from outside. From inside, both looked like rigour, right up until someone who could plainly speak the language walked out with a fail.
The lesson we keep is bigger than one list of questions. A pass mark only means something if people can fail it. And they have to be able to fail it for the right reasons.
