AI in Learning and Development: Field Notes

What we have learned using AI to train and assess people at work, as a dated chronology of our own experiments. What we tried, what the numbers said, what we changed, and what failed.

We ask buyers to demand evidence from the vendors they evaluate, including us. This is where ours is kept, with the sample sizes and the limits attached.

Newest first. If you are starting here, start with Field Note 01, can AI assess soft skills? It is the method the rest of these rest on.

Field Note 04  /  4 September 2026

The obvious fix was half wrong

Our grader was worst at two questions, and both asked two things at once. We split them. It was right for one and wrong for the other, and then a larger check reversed what we thought we had learned about why.

The two compound questions score 0.63 and 0.75 when a human is checked against himself. The clearest single-idea question scores 0.97. 36 conversations, one rater.

Field Note 03  /  26 August 2026

How to write an AI rubric that does not fail good people

Our AI certification failed people who were good at the job. The cause was our own list of questions, which rewarded describing the skill instead of doing it. Three rules for writing a rubric that does not.

5 of 13 questions could not be passed by doing the job well. After the fix, a strong performer scored 90. The pass mark never moved.

Field Note 02  /  26 August 2026

Does a bigger AI model score people better? We tested six.

We ran six AI scorers against a human rater on the same 36 conversations. The flagship agreed with the human least, then nine weeks later the ranking reversed. The result is real, and the lesson most people take from it is the wrong one.

June: cheapest two kappa 0.87, flagship 0.84. August, same corpus: the model that came last now comes first.

Field Note 01  /  26 August 2026

Can AI assess soft skills? Ours failed first. Here is the fix.

A competent conversation scored 50. The same transcript, scored twice, returned two different numbers. What we changed, and why it is the reason any number we publish is worth reading.

ICC 0.955 and weighted kappa 0.857 against blind human rating. 36 transcripts, one skill, one rater.