Miloš Kankaraš | 2026 | Preprint
Public leaderboards order AI models by a point score, and the debate about them turns on leads of one or two points. This paper asks how many of those leads the benchmarks can detect, using the item-level results that 61 public leaderboards publish. 55 of the 58 leaderboards that discriminate at all cannot fully separate their own top five models, and on GPQA, SWE-bench Verified and tau2-bench the headline comparisons between frontier models are statistical ties.
Full text: doi:10.5281/zenodo.23036392
Data and code: doi:10.5281/zenodo.22998207
The study in brief, with charts
Suggested citation
Miloš Kankaraš (2026). Which AI model is truly the best: how reliable are public benchmarks of frontier AI models? Zenodo (preprint). 10.5281/zenodo.23036392