Calibration is an Operations Problem
How a measured result becomes an institutional claim, and what an editor can do to make that transition visible.
About
Much of my work has involved checking whether a document says what the evidence supports.
I've enforced journal editorial standards across researchers who didn't report to me. At a startup, I checked the marketing against a product that changed faster than the copy. I've also written product documentation and ghostwritten for founders. A sentence can be clear and still be wrong. Finding that out before publication was part of the work.
On Rougarou, I bring that experience to AI evaluations. I'm interested in how a benchmark was constructed and how a company gets from a result to a claim about safety. A good footnote can do more for a document's credibility than another page of assurances.
Strong opinions held loosely means being willing to say what I think, then show what would change my mind. If I've misread a source, I want to hear about it.
How a measured result becomes an institutional claim, and what an editor can do to make that transition visible.
A GPT-5 footnote explains why two results on the same benchmark differ. That small disclosure should be standard practice.
The same models beat one expert baseline and miss another. The choice of comparison can write the headline before anyone reads the results.
What 380 hours of adversarial testing can tell us about a model's safeguards, and why the definition of a finding changes the arithmetic.