The two worst questions
Our certification does not give a conversation one overall grade. It asks a series of narrow yes-or-no questions about specific moments, and the arithmetic does the rest. One of those questions asked whether the manager stayed composed under pushback, which in practice meant two separate things: that they did not cave, and that they did not get aggressive. Another asked whether they acknowledged the other person's point and then brought the responsibility back to them.
Both are compound. Every market researcher already knows why that is dangerous, because it is the oldest flaw in survey design: never ask "was the service fast and friendly?" When someone answers no, you cannot tell which half they meant, and when they hesitate you cannot tell which half they are hesitating over. The technical name is a double-barrelled question, and we had two of them sitting in a rubric that decides whether somebody earns a credential.
They were also, measurably, our two weakest questions. So the hypothesis wrote itself: split them, and the problem goes away.
Splitting the first one worked beautifully
We separated "did not cave" from "did not get aggressive" and re-scored every conversation. The two halves turned out to be wildly unequal. Almost nobody gets aggressive with a colleague in a practice roleplay, so that half was true nearly all the time and carried no information at all. The other half, whether the manager held their position, was where the real difference between people lived.
Bolted together under a single "stayed composed" heading, the easy half was dragging the hard one around. Worse, the compound phrasing read to the grader as a condition where everything had to be plainly true at once, so it withheld marks whenever either half was unclear. That is where a chunk of our grader's excess strictness had been hiding.
Split apart, the strictness on that question vanished entirely for one model, and fell for the other.
Splitting the second one made it worse
We ran exactly the same procedure on the other question and it backfired. Every way we tried to put the halves back together produced a worse result than leaving it alone, and the loosest recombination was three times worse.
The reason is worth more than the result. Acknowledging someone's point and then re-anchoring responsibility is not two things a person does. It is one thing, and doing only the first half is not partial credit, it is a different behaviour entirely. It is the manager who nods sympathetically and then lets the issue go. Splitting that question let the easy half buy credit for a conversation that never did the hard half, which is precisely the failure the question exists to catch.
So the rule we came away with, and now design by: one question, one idea. If a rubric line has an "and" in it, work out whether you are looking at two separate things that got stapled together, or one thing that genuinely needs two clauses to describe. Split the first. Never split the second.
The reversal, and then the reversal of the reversal
Here is where we nearly got it wrong in public.
Two months after the split experiment, I re-rated a sample of the same conversations blind, to see whether a human found those two questions hard. I did not. On that sample I agreed with my own earlier marking perfectly on both. That is a genuinely different diagnosis. It says the questions were clear enough for a person and the machine was the weak link, which would mean the fix was a better grader rather than a better rubric.
We held off on that conclusion, because the sample was nine conversations and five of them were so clearly strong or clearly weak that re-rating them was close to automatic.
This week I finished the other twenty seven, so the check now covers every conversation in the study rather than the easy quarter of it. It reversed the reversal. Across all thirty six, the two double-barrelled questions are the ones I disagree with myself about most, and the single-idea questions are the ones I reproduce almost perfectly. The earlier result was not wrong so much as it was measured on a sample too easy to reveal anything.
The wording was the problem after all. And the evidence is better than we expected, because the difficulty shows up in the same two places for a person and for a machine. A question that is ambiguous is ambiguous for everybody.
What to do with this
If you are evaluating anyone's assessment, ours included, the question underneath all of this is not "is your AI accurate". It is "how reliable is each line of your rubric, and how do you know?"
Three things you can ask for, and they are all reasonable.
Show me your rubric lines. Count the "ands". Every compound line is a place where two markers can disagree while both being right, and where a score can move without anybody's performance changing.
Show me agreement per line, not just overall. An overall figure averages your worst line together with your best and hides exactly the thing you need to see. Ours ranges from 0.97 down to 0.63 on the same instrument.
Ask whether a human was ever checked against himself. Not against the machine, against his own earlier marking, blind, after enough time that he cannot remember. It is the cheapest measurement in assessment and almost nobody does it, which is a shame, because until you have it you have no idea whether a disagreement between a human and a machine means the machine is wrong or the question is.
We got two of these three wrong before we got them right, and the only reason we know is that we kept measuring after we had an answer we liked.
For how the scoring instrument was built in the first place, see the score that wandered. For why the most capable model was the worst scorer, see method beats model size. For the time our own checklist rewarded describing the skill over doing it, see the narration trap.
One rater, 36 conversations, one skill. A question that is ambiguous is ambiguous for everybody.
