The Aggregate: unified LLM rankings, daily
We fold every public benchmark we can find into one daily ranking. Unified rankings for 3078 AI models in the fused corpus across 10227 public benchmarks in the current fit, aggregated from public leaderboards and refreshed daily. Rankings use an Item Response Theory (IRT) fit that maps every model onto one ELO scale.
Top models by unified ELO
| Rank | Model | Provider | Unified ELO | Benchmarks |
|---|---|---|---|---|
| 1 | Claude Mythos 5.1 | Anthropic | 2163 ± 29 | 21 |
| 2 | Claude Mythos 5 | Anthropic | 2143 ± 13 | 104 |
| 3 | GPT-6 | OpenAI | 2121 ± 7 | 513 |
| 4 | GPT-6 Pro | OpenAI | 2120 ± 27 | 16 |
| 5 | Claude Mythos Preview | Anthropic | 2115 ± 13 | 96 |
| 6 | Claude Fable 5.1 | Anthropic | 2114 ± 7 | 440 |
| 7 | Gemini 3.8 Flash Cyber | 2101 ± 33 | 14 | |
| 8 | Claude Opus 5 | Anthropic | 2077 ± 5 | 895 |
| 9 | Claude Fable 5 | Anthropic | 2051 ± 5 | 818 |
| 10 | GPT-5.5 Pro | OpenAI | 2040 ± 15 | 69 |
| 11 | GPT-5.6 Sol | OpenAI | 2031 ± 4 | 1251 |
| 12 | Gemini 3.8 Flash | 2031 ± 8 | 412 | |
| 13 | GPT-5.6 Pro Sol | OpenAI | 2015 ± 18 | 55 |
| 14 | Gemini 3.7 Flash | 2010 ± 7 | 530 | |
| 15 | Muse Spark 1.3 | Meta | 2000 ± 9 | 292 |
| 16 | Grok 4.6 | xAI | 1993 ± 7 | 445 |
| 17 | GPT-5.5 | OpenAI | 1988 ± 3 | 1865 |
| 18 | Kimi K3 | Moonshot | 1983 ± 5 | 882 |
| 19 | GPT-5.4 Pro | OpenAI | 1983 ± 13 | 79 |
| 20 | Qwen 3.8 Max (0902) | Alibaba | 1975 ± 19 | 97 |
Which AI is worth paying for?
Premium subscriptions to budget routers. One overall value score. Editorial tier list.
| Plan | Tier | Score / 100 | Price | Verdict |
|---|---|---|---|---|
| ChatGPT Plus | A | 84 | USD 20/mo | A strong all-round bundle; sustained agent work can outgrow its limits. |
| ChatGPT Pro 5x | A | 84 | USD 100/mo | More room for serious chat and coding work without the largest monthly bill. |
| ChatGPT Pro 20x | A | 83 | USD 200/mo | A large chat-and-coding allowance that pays off when you actually use it — but OpenAI is not currently selling it to new subscribers. |
| Google AI Pro | A | 82 | USD 19.99/mo | Broad Gemini and Google-app benefits at an accessible monthly price. |
| Claude Pro | A | 82 | USD 20/mo | Strong chat and coding in one plan; both draw from the same variable allowance. |
| GitHub Copilot Pro+ | A | 82 | USD 39/mo | Premium model choice and a larger published credit pool for regular coding. |
| Google AI Ultra 5x | A | 82 | USD 99.99/mo | Higher Gemini app and Antigravity limits plus creative credits for frequent users of the Google bundle. |
| GitHub Copilot Max | A | 82 | USD 100/mo | A sizeable published coding budget for people who will use the extra credits. |
Latest leaderboard changes
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Recall (94.28 vs 86.55 by Troml)
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench (86.58 vs 80.34 by metor.com)
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Correctness (89.8 vs 83.8 by Troml)
- LimiX-2 (default) took the lead on TabArena All Tasks (93.1 vs 87.3 by TabFM (default))
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Completeness (90.62 vs 86.22 by metor.com)
- Muse-Glimmer-30B (high, 16k) took the lead on ReasonScape R12 (954.4 vs 951.64 by Qwen3.5-397B-A17B (AWQ, 16k) (Thinking))
- Mixedbread (+ Opus 5) took the lead on EnterpriseRAG Bench - Valid Extra Docs Resistance (99.61 vs 99.53 by OpenClaw)
Explore: all models · benchmark catalog · capability trends · daily changes
Interactive version: theaggregate.ai/ · How It Works · Data refreshed daily, snapshot 2026-09-19.