Every few weeks a new AI model launches with a chart. The chart has bars, the bars have percentages, and the percentages are supposed to tell you which model is smarter. Here is the uncomfortable thing worth knowing before you trust the next chart: a growing body of research is now explicitly arguing that these benchmarks often don't measure what they claim to.
Some of it is gaming, models trained in ways that happen to be very good at the specific test without the underlying skill improving nearly as much. Some of it is simpler: a benchmark built to measure one thing (raw capability) gets quietly used to imply something else entirely (safety, or reliability, or "better at your actual job"), and nobody flags the switch.
None of this means ignore every number. It means treat a benchmark score the way you'd treat a single stranger's five-star review: a data point, not a verdict. The people actually building with these tools day to day are increasingly saying the same thing out loud: cross-check the leaderboard against your own small test, on your own actual task, before you believe the marketing chart.