Live data
Assembling the frontier
Ranking the latest models and provider endpoints.
Loading model dataLive data
Ranking the latest models and provider endpoints.
Loading model dataAgentic model routing
Using Artificial Analysis API benchmark data, cached for up to 24 hours. The source measurement date is not supplied. Sources and refresh policy.
Agentic fit lab
Set your workload, minimum context, and ranking preferences.
Task
Agentic Index preference
Models below this target rank lower. It does not certify autonomy or reliability.
Minimum context
Optimize for
Shared links preserve these settings and exact model variants. Recommendations are recalculated from current data and may change.
Fit is a WhatLLM screening score, not a benchmark or a probability of success. Every result has an observed Agentic Index, meets the minimum advertised context window, and respects the open weights filter when selected. Context capacity does not measure retrieval accuracy.
For codebase agent, the base weights are Agentic Index 34%, Intelligence Index 12%, Coding Index 22%, terminal benchmarks 16%, τ-Bench 4%, GDPval 4%, context 4%, cost 2%, and response time 2%.
Missing metrics earn no ranking credit. Index and percentage scores are scaled to 0–1; GDPval Elo uses a WhatLLM 900–1800 scale. Cost and time are normalized across the current dataset. Your optimization preference adds weight to the selected measure, and models below your Agentic Index target receive a penalty.
Cost per task means an Artificial Analysis Intelligence Index evaluation task. It is not cost per successful agent task. Check the exact model variant, provider, and results on your own workload before choosing. Sources: Artificial Analysis and WhatLLM methodology.
Fit
65
Agentic
50.5
Cost / task
$1.99
Response
2.0m
Context
1.0M
Start your evaluation here: GPT-5.6 Sol (max) leads this workload’s weighted ranking, with an observed Agentic Index of 50.5 and 1.0M context. It is below your preferred Agentic Index target.
AA evaluation task cost exceeds $1.50; check costs on your workload.
Check this model’s pricing and providersValue route
OpenAI
Fit
39
Task cost
$0.085
Response
51.8s
Benchmark cost and response time describe AA evaluations; validate your own agent loop.
Fast route
OpenAI
Fit
59
Task cost
$0.808
Response
23.7s
Benchmark cost and response time describe AA evaluations; validate your own agent loop.
Open route
Alibaba
Fit
58
Task cost
$0.372
Response
50.6s
Benchmark cost and response time describe AA evaluations; validate your own agent loop.
Frontier map
11 meet the Agentic Index target
131 with cost/task
GPT-5.6 Sol (max)
OpenAI
Agentic
50.5
Task cost
$1.99
| Model | Fit | Agentic | Task cost | Response | Context |
|---|---|---|---|---|---|
| GPT-5.6 Sol (max) OpenAI · Proprietary | 65 | 50.5 | $1.99 | 2.0m | 1.0M |
| GPT-5.6 Sol (xhigh) OpenAI · Proprietary | 63 | 47.8 | $1.18 | 49.3s | 1.0M |
| Claude Fable 5 (Max Effort, Opus 4.8 Fallback) Anthropic · Proprietary | 59 | 51 | $8.75 | 1.5m | 1.0M |
| GPT-5.6 Sol (high) OpenAI · Proprietary | 59 | 44.7 | $0.808 | 23.7s | 1.0M |
| Muse Spark 1.3 (max) Meta · Proprietary | 59 | 55.7 | $1.60 | 33.8s | 1.0M |
| Qwen3.8-Flash-Next Alibaba · Open | 58 | 53.9 | $0.372 | 50.6s | 256K |
| Muse Spark 1.3 (xhigh) Meta · Proprietary | 58 | 51.8 | $1.37 | 31.9s | 1.0M |
| Grok 4.6 (high) SpaceXAI · Proprietary | 58 | 53.4 | $1.86 | 37.7s | 500K |
| GLM-5.3 (max) Z AI · Open | 57 | 53.4 | $2.01 | 41.0s | 1.0M |
| GLM-5.3-Flash Z AI · Open | 57 | 51.2 | $0.253 | 26.6s | 1.0M |
A small per-token price gap can become a large bill when an agent runs many turns, calls tools, and carries long state. Cost per Intelligence Index task adds evaluation cost to the comparison. It does not estimate the cost of your complete agent workflow or guarantee a successful task.
Compare the recommended model with a cheaper candidate on representative tasks. Record successful completions, retries, tool calls, total cost, and elapsed time before selecting a model for production.
Ranking library
Focused rankings for the decisions engineers actually make.