The Mascot Test: a walrus playing a cello

Most AI benchmarks are numbers. A model scores 88.7 on some acronym, another scores 91.2 — and unless you work in the field, neither number tells you anything about what these models can actually do. This is the first task in a series of human-friendly benchmarks: small, real tasks where you don’t need a leaderboard to understand the result — you can just look at it. The task Every model gets the exact same prompt, once, with no retries and no cherry-picking: ...

July 19, 2026 · 4 min · Denis Ilguzin