Command Palette
Search for a command to run...
Series
Flaky by Default
Evaluating AI systems, written by someone who spent 13 years in test automation before the systems started answering back. Benchmarks rot once they are worth gaming. Every instrument has a noise floor and you should know what yours is. A pass rate that changes between runs is a number with a confidence interval, not a number. Anything you optimise stops being a measure. Posts on benchmarks, LLM-as-judge, contamination, nondeterminism, eval harness design and agent reliability: the measurement layer underneath the hype.

No posts yet

