Field Notes

A dated chronology of our own experiments and their results. What we tried, what the numbers said, what we changed, and what failed.

We ask buyers to demand evidence from the vendors they evaluate, including us. This is where ours is kept, with the sample sizes and the limits attached. Entries are dated and never rewritten after publication. When the numbers move, we say so and the original stays up.

Field Note 01  /  26 August 2026

The score that wandered

A competent conversation scored 50. The same transcript, scored twice, returned two different numbers. What we changed, and why it is the reason any number we publish is worth reading.

ICC 0.96 and weighted kappa 0.87 against blind human rating. 36 transcripts, one skill, one rater.

Field Note 02  /  26 August 2026

Method beats model size

We ran six AI scorers against a human rater on the same 36 conversations. The flagship agreed with the human least. The result is real, and the lesson most people take from it is the wrong one.

Cheapest two: kappa 0.87. Flagship: kappa 0.84, and nearly twice as harsh.

Field Note 03  /  26 August 2026

The narration trap

For weeks our certification failed people who were good at the job. The cause was not the bar. It was our own checklist, quietly asking learners to describe the skill instead of do it, and every defence we had built waved it through.

5 of 13 checks were unpassable by doing the skill. After the fix, a natural run scored 90. The bar never moved.