Skip to content

Study

Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

Most of the leads that make headlines on AI leaderboards are smaller than the benchmarks can measure. This study checks 61 public leaderboards against their own published results.

When a new AI model is released, it arrives with a table of benchmark scores, and the discussion that follows is about who leads by a point or two. A benchmark is a test, and like any test it measures with some error. This study asks a simple question: how many of those leads are larger than that error?

We checked 61 public leaderboards, from GPQA and SWE-bench to tau2-bench, using only the question-by-question results their own maintainers publish. The answer is that most of the leads people argue about are too small for the benchmarks to detect.

The leaders are often tied

On GPQA, a test of graduate-level science, Gemini 3 Pro scores 80.3 per cent and GPT-5 scores 79.1. The difference is well inside the margin of error: the two are statistically tied. So are several other pairs near the top.

Chart: the eight highest scores on GPQA with their margins of error. Gemini 3 Pro (80.3 per cent) and GPT-5 (79.1 per cent) overlap and are statistically tied.

The newest models are no easier to rank

On tau2-bench, where models handle customer-service conversations, the newest models are Claude Opus 5, GPT-5.6, Grok 4.5 and Qwen 3.8 Max. None of the comparisons between them that we tested is a real difference. On SWE-bench, a widely used coding benchmark, Claude Opus 4.6 cannot be told apart from the earlier Opus 4.5.

Chart: eleven of the newest models on tau2-bench banking tasks with their margins of error. Qwen 3.8 Max, Claude Opus 5 and Grok 4.5 are statistically tied at the top.

How big a lead has to be

The fewer questions a benchmark has, the bigger a lead must be before it means anything. On GPQA a lead needs about 4 to 5 points; on tau2-bench, with 97 tasks, 7 to 9. Many of the leads reported at a launch are one or two points. Very large tests such as MMLU can detect small differences, but the leading models have outgrown them.

Chart: the lead each benchmark needs before a difference is real, from 0.4 points on MMLU to between 7 and 9 points on tau2-bench.

The test setup can matter more than the model

The same model, Claude Opus 4.7, was run twice on tau2-bench by the benchmark’s own team, on two versions of the testing software. It scored 25 per cent once and 40 per cent the other time. The public leaderboard lists results from both versions side by side, so part of what looks like progress between models is a change in the test.

Chart: Claude Opus 4.7 scores 25.3 per cent on an earlier version of the test harness and 40.2 per cent on the current version.

A pattern across the field

Across all the leaderboards, 55 of 58 cannot fully tell apart their own top five models, and 16 cannot separate any of them. The rankings are reliable for models far apart and unreliable exactly where attention is focused, at the top.

Chart: one square per leaderboard. Only 3 of 58 leaderboards fully separate their top five models; 16 separate none of them.

What this means

A lead that a benchmark cannot detect is not evidence that one model is better. Leaderboards would be more honest if they showed margins of error and marked ties, used tests large enough for the differences they are asked to judge, and compared only models tested the same way.

The paper and the data

Which AI model is truly the best: how reliable are public benchmarks of frontier AI models? Miloš Kankaraš, 2026. Preprint, doi:10.5281/zenodo.23036392.

Zenodo: report cards for all 61 leaderboards, head-to-head results and analysis code, doi:10.5281/zenodo.22998207

The analysis uses only results published by the benchmarks’ maintainers; no model was re-run. Results reflect the model versions each benchmark tested, as of September 2026.

Part of Measurement comparability

Related studies

Published as

  1. 2026Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?Zenodo (preprint) DOIPreprint

Research programmes Every study Publications