The definitive ranking of AI models for software development, code generation, and programming. Ranked by LiveCodeBench, Terminal-Bench, and SciCode โ independent evaluations refreshed through the live data layer.
Quality Index
63.1
Anthropic
Quality Index
62.1
Anthropic
Quality Index
60.9
OpenAI
| Rank | Model | Quality | LiveCodeBench | Terminal-Bench | SciCode | License |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic | 63.1 | - | - | 56% | Proprietary |
| 2 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic | 62.1 | - | 63% | 60% | Proprietary |
| 3 | GPT-5.6 Sol (max) OpenAI | 60.9 | - | 66% | 56% | Proprietary |
| 4 | Grok 4.6 (high) SpaceXAI | 60.9 | - | - | 54% | Proprietary |
| 5 | Kimi K3 (max) Kimi | 59.7 | - | - | 59% | Open |
| 6 | GLM-5.3 (max) Z AI | 59.5 | - | - | 56% | Proprietary |
| 7 | Qwen3.8 Max Alibaba | 58.1 | - | - | 53% | Proprietary |
| 8 | Qwen3.8 2.4T A95B Alibaba | 57.7 | - | - | 52% | Open |
| 9 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic | 57.3 | - | 58% | 54% | Proprietary |
| 10 | Muse Spark 1.2 (xhigh) Meta | 56.8 | - | - | 56% | Proprietary |
Rankings combine the Artificial Analysis Coding Index with task-level benchmarks that test different parts of programming work:
Fresh code-generation problems across languages. Useful for raw coding ability, but not a complete repository-agent test.
Tests complex terminal operations, shell scripting, DevOps work, and recovery over longer trajectories.
Scientific computing and research programming. Tests ability to implement algorithms from papers and numerical methods correctly.
Quality Index from Artificial Analysis. Read which coding benchmark matters before treating one score as a universal ranking.
See exact pricing, latency, and benchmark scores for all 10 coding models in our interactive comparison tool.
Claude Opus 5 (Adaptive Reasoning, Max Effort) currently leads this multi-benchmark coding ranking. The right model can still change for repository repair, terminal work, scientific programming, or local deployment.
Start with the top Coding Index and Terminal-Bench models above, then run a private evaluation on recent issues from your repositories. Public rankings do not measure your agent harness, review standards, or failure cost.
Kimi K3 (max) is the current open-weight leader in this dataset with a Coding Index of 76.2. Verify its license, total parameter count, and provider or hardware fit before deployment.
Qwen3-Coder 30B is the practical 24GB-class option, while Qwen3-Coder-Next targets systems with at least 64GB of usable memory. See the hardware-tier guide for artifact sizes and commands.
Compare the current OpenAI and Anthropic variants in the live table rather than relying on a fixed brand verdict. Use the same harness for repository work and include accepted-result cost, latency, and review quality.
The page loads Artificial Analysis data at runtime when configured and falls back to the checked-in snapshot during build or API failure. Editorial copy and methodology were reviewed on July 16, 2026.
Self-hostable models
๐งฎAIME 2025 rankings
๐คTool use & agents
โกAny 2โ4 models
Data sources: Rankings based on the Artificial Analysis Intelligence Index. Explore all models in our interactive leaderboard or compare models side by side.