benchmarks
Every record this desk has filed under benchmarks, newest first, each with the number of sources it can still show you.
Opinion: when the labs grade their own homework, the receipts desk beats the leaderboard.
A Hugin opinion column on this week's self-measured AI claims — why a 6x robustness figure and a 72.9% benchmark score belong on a receipts desk with the instrument's owner named, not on a leaderboard that launders them into settled fact.
Also filed underopinionai-safetyevidence-postureeditorial-disciplinesource-receipts
A record appears here because it carries benchmarks in its own frontmatter. If a record you expected is missing, it was filed under a different subject — the full list is on the topics index.