Calibration is an Operations Problem
How a measured result becomes an institutional claim, and what an editor can do to make that transition visible.
Beau Henry
I write about what a technical document lets you claim, and what it leaves unresolved. These essays read AI evaluations closely enough to argue with them.
About me and my work →How a measured result becomes an institutional claim, and what an editor can do to make that transition visible.
A GPT-5 footnote explains why two results on the same benchmark differ. That small disclosure should be standard practice.
The same models beat one expert baseline and miss another. The choice of comparison can write the headline before anyone reads the results.
What 380 hours of adversarial testing can tell us about a model's safeguards, and why the definition of a finding changes the arithmetic.
The archive includes everything I've published here, along with a guest essay by Jake Gaylor.
Browse the full archive →