The Aggregate — unified LLM rankings, daily
Unified rankings for 1776 AI models across 5206 public benchmarks, aggregated from public leaderboards and refreshed daily. Rankings use an Item Response Theory (IRT) fit that maps every model onto one ELO scale.
Top models by unified ELO
| Rank | Model | Provider | Unified ELO | Benchmarks |
|---|---|---|---|---|
| 1 | Human Expert | Human | 2500 ± 48 | 59 |
| 2 | GPT-5.5 Pro (xHigh) | OpenAI | 2268 ± 57 | 12 |
| 3 | Claude Mythos 5 | Anthropic | 2250 ± 31 | 57 |
| 4 | Gemini 3 Deep Think | 2186 ± 71 | 18 | |
| 5 | GPT-5.6 Sol (Max) | OpenAI | 2175 ± 48 | 38 |
| 6 | Claude Fable 5 (Max) | Anthropic | 2173 ± 37 | 27 |
| 7 | Claude Opus 4.6 (Medium) | Anthropic | 2129 ± 131 | 7 |
| 8 | Claude Fable 5 (High) | Anthropic | 2104 ± 41 | 18 |
| 9 | Claude Mythos Preview | Anthropic | 2103 ± 21 | 79 |
| 10 | GPT-5.6 Sol (xHigh) | OpenAI | 2096 ± 39 | 42 |
| 11 | GPT-5.5 Pro | OpenAI | 2092 ± 63 | 29 |
| 12 | Claude Fable 5 | Anthropic | 2088 ± 25 | 140 |
| 13 | GPT-5.5 (xHigh) | OpenAI | 2057 ± 17 | 109 |
| 14 | GPT-5.6 Terra (Max) | OpenAI | 2051 ± 39 | 31 |
| 15 | GPT-5.6 Sol (High) | OpenAI | 2041 ± 33 | 28 |
| 16 | Kimi K3 (Max) | Moonshot | 2034 ± 52 | 11 |
| 17 | GPT-5.4 Pro (xHigh) | OpenAI | 2019 ± 27 | 63 |
| 18 | GPT-5.6 Sol (Medium) | OpenAI | 2016 ± 31 | 23 |
| 19 | GPT-5.2 (xHigh) | OpenAI | 2012 ± 26 | 85 |
| 20 | GPT-5.5 (High) | OpenAI | 2011 ± 26 | 66 |
Latest leaderboard changes
- Gemini 3.6 Flash took the lead on RuneBench (7454 vs 7400 by GPT-5.6 Sol)
- Claude Fable 5 took the lead on HiL-Bench (56.33 vs 29.1 by GPT-5.5)
- Troml took the lead on EnterpriseRAG Bench - Completeness (81.84 vs 72.86 by OpenClaw)
- Troml took the lead on EnterpriseRAG Bench (76.79 vs 68.22 by OpenClaw)
- Troml took the lead on EnterpriseRAG Bench - Recall (86.55 vs 79.02 by OpenClaw)
- Grok 4.5 took the lead on Vals AI SkillsBench (66.03 vs 62.55 by GPT-5.5 Codex)
- Troml took the lead on EnterpriseRAG Bench - Correctness (83.8 vs 81.6 by OpenClaw)
- Grok 4.5 (xHigh) took the lead on BoxPwnr CTF Bench (56.73 vs 55.47 by GLM-5.1)
Explore: all models · benchmark catalog · capability trends · daily changes
Interactive version: theaggregate.ai/ · How the rankings work · Data refreshed daily, snapshot 2026-07-22.