Skip to content

Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

Miloš Kankaraš | 2026 | Preprint

Public leaderboards order AI models by a point score, and the debate about them turns on leads of one or two points. This paper asks how many of those leads the benchmarks can detect, using the item-level results that 61 public leaderboards publish. 55 of the 58 leaderboards that discriminate at all cannot fully separate their own top five models, and on GPQA, SWE-bench Verified and tau2-bench the headline comparisons between frontier models are statistical ties.

Full text: doi:10.5281/zenodo.23036392

Data and code: doi:10.5281/zenodo.22998207

The study in brief, with charts

Suggested citation

Miloš Kankaraš (2026). Which AI model is truly the best: how reliable are public benchmarks of frontier AI models? Zenodo (preprint). 10.5281/zenodo.23036392


Related