Live data
Assembling the frontier
Ranking the latest models and provider endpoints.
Loading model dataLive data
Ranking the latest models and provider endpoints.
Loading model dataLive model ledger
Need methodology and scenario-specific picks? Read the dedicated ai models for ollama guide.
Top overall
Kimi K3
QI 59.7
Best value
DeepSeek V4 Flash 0731
$0.05/M
Fastest
Celeris-1
2040 tok/s
Largest context
Llama 4 Scout
1.3M
| # | Model | Quality | Price/M | Speed | Context |
|---|---|---|---|---|---|
| 1 | Kimi K3Kimi | 59.7 | $5.1 | 167 tok/s | 1.0M |
| 2 | GLM-5.3Z AI | 59.5 | $2.1 | 66 tok/s | 1.0M |
| 3 | Qwen3.8 MaxAlibaba | 58.1 | $3.0 | 26 tok/s | 1.0M |
| 4 | Qwen3.8 2.4T A95BAlibaba | 57.7 | $2.8 | 112 tok/s | 984K |
| 5 | GLM-5.3-FlashZ AI | 57.5 | $0.12 | 297 tok/s | 1.0M |
| 6 | DeepSeek V4 Pro 0813DeepSeek | 53.2 | $1.6 | 166 tok/s | 1.0M |
| 7 | GLM-5.2Z AI | 52.6 | $0.92 | 354 tok/s | 1.0M |
| 8 | Qwen3.8 27B (xhigh)Alibaba | 52 | $1.1 | 47 tok/s | 1.0M |
| 9 | DeepSeek V4 Flash 0731DeepSeek | 51.8 | $0.05 | 328 tok/s | 1.0M |
| 10 | DeepSeek V4 Flash VisionDeepSeek | 51.5 | $0.66 | 117 tok/s | 1.0M |
| 11 | Kimi K3 (low)Kimi | 48.3 | $6.0 | 96 tok/s | 1.0M |
| 12 | MiniMax-M3MiniMax | 45.4 | $0.41 | 228 tok/s | 1.0M |
| 13 | DeepSeek V4 ProDeepSeek | 45.3 | $0.54 | 162 tok/s | 1.0M |
| 14 | Qwen3.8 27B (medium)Alibaba | 44.5 | $1.1 | 51 tok/s | 1.0M |
| 15 | Kimi K2.7 CodeKimi | 43 | $1.4 | 322 tok/s | 262K |
| 16 | MiMo-V2.5-ProXiaomi | 42.9 | $0.54 | 80 tok/s | 1.0M |
| 17 | Qwen3.8 27B (low)Alibaba | 42.9 | $1.1 | 51 tok/s | 1.0M |
| 18 | Inkling (xhigh)Thinking Machines | 42.3 | $1.7 | 242 tok/s | 524K |
| 19 | Hy3Tencent | 42.2 | $0.23 | 74 tok/s | 262K |
| 20 | Nex-N2-ProNex AGI | 41.7 | $1.0 | 139 tok/s | 262K |
| 21 | Inkling SmallThinking Machines | 41.2 | $0.52 | 344 tok/s | 524K |
| 22 | Agnes 2.5 Pro AlphaSapiens AI | 39.7 | $0.56 | 132 tok/s | 1.0M |
| 23 | Qwen3.7 PlusAlibaba | 39.4 | $0.70 | 56 tok/s | 1.0M |
| 24 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 38.3 | $0.93 | 339 tok/s | 512K |
| 25 | MiMo-V2.5Xiaomi | 38 | $0.18 | 68 tok/s | 1.0M |
| 26 | Ling 3.0 FlashInclusionAI | 37.8 | $0.09 | 406 tok/s | 262K |
| 27 | Qwen3.6 27B (Reasoning)Alibaba | 37.7 | $0.34 | 447 tok/s | 262K |
| 28 | Muse Glimmer (high)Meta | 35.1 | $0.52 | 130 tok/s | 131K |
| 29 | Qwen3.8 27BAlibaba | 34.7 | $1.1 | 51 tok/s | 1.0M |
| 30 | Qwen3.5 397B A17B (Reasoning)Alibaba | 34.3 | $0.88 | 134 tok/s | 262K |
Showing top 30 of 114 ranked models
View all in Explore →Decision routes
The overall score opens the conversation. These guides finish it with task-specific methodology and picks.
Ranking library
Focused rankings for the decisions engineers actually make.
Methodology
Every model is scored using the Artificial Analysis Intelligence Index — a composite of GPQA Diamond, AIME 2025, LiveCodeBench, MMLU-Pro, and other benchmarks, weighted into a single 0-100 quality score. Speed, price, and context window are tracked live across providers.
The overall ranking is a starting point. For production decisions, narrow by use case using the category pages above, then compare finalists head-to-head on Compare.
Field notes
Kimi K3 leads on overall quality right now, but the best model depends on your priorities. Coding, cost, speed, and context length all shift the answer. Use the category rankings above to find the right fit.
DeepSeek V4 Flash 0731 currently offers one of the best quality-to-cost ratios. Open-source models on providers like Groq or Together can be even cheaper at strong quality levels.
Start with overall quality index, then narrow by what matters for your workload: cost per million tokens, output speed, context window, or a specific capability like coding or tool use. Use our Compare tool to put finalists head to head.
Celeris-1 leads on output speed right now at 2040 tokens/second. Speed matters most for real-time applications and agentic workflows with many sequential steps.
Llama 4 Scout has the biggest context window in this ranking at 1.3M. For a dedicated long-context comparison, see our largest context window page.
Data is pulled from Artificial Analysis and refreshed automatically. New models appear as soon as they have benchmark scores and provider endpoints. The ranking reflects the live state of the leaderboard.