A benchmark gives you a score. I build the thing that tells you whether the score is real — simulators, oracles and audits that can say no.
Each of these re-runs its own experiment in the page, next to the original implementation, from the same data the study used. Neither is a recording of a result.
Two identical fleets consume a year of real Seattle medic calls at the same clock. The engine checks itself against the Python reference on load and currently agrees to 0.05 seconds.
OPEN INSTRUMENTEnumerate every combination path for any puzzle you type, then watch two optimisers run the same reward until one has a gradient of exactly zero and can never move again.
OPEN INSTRUMENTMost of what I build is a way of checking something I already believe. The EMS study started as "dynamic ambulance relocation helps, let's measure how much" and ended with the finding that parking the same units somewhere better beats moving them all year. The self-play study started as "a generator will learn to sit at the solver's edge" and ended with three separate retractions of my own claims, each one forced by a test I wrote to try to break the last one.
So the instruments matter more than the results. An exact oracle, a difficulty-matched control, a regression suite that fails when a headline stops being true — those are the parts that survive.