Agentic model routing
Match a workload to models using Agentic Index, task-level cost, response time, benchmark signals, and context requirements. The shortlist is built from the same live Artificial Analysis data used across WhatLLM.
Agentic fit lab
Ranking changes as task shape, autonomy, context, and economics change.
Task
Autonomy needed
Context load
Optimize for
Best fit
OpenAI
Fit
92
Agentic
51.8
Cost / task
$0.682
Response
1.0m
Context
1.0M
Value route
OpenAI
Fit
62
Task cost
$0.095
Response
10.5s
Good balance across score and economics
Fast route
OpenAI
Fit
88
Task cost
$0.453
Response
20.2s
Good balance across score and economics
Open route
Z AI
Fit
75
Task cost
$0.468
Response
13.8s
Good balance across score and economics
Frontier map
14 high-fit models
112 with cost/task
GPT-5.6 Sol (xhigh)
OpenAI
Agentic
51.8
Task cost
$0.682
Shortlist
GPT-5.6 Sol (xhigh)
51.8 agentic score
92
fit
GPT-5.6 Sol (max)
54 agentic score
89
fit
GPT-5.6 Sol (high)
48.5 agentic score
88
fit
Claude Fable 5 (Max Effort, Opus 4.8 Fallback)
52.8 agentic score
84
fit
Kimi K3
50.1 agentic score
82
fit
GPT-5.6 Sol (medium)
44.5 agentic score
80
fit
GPT-5.6 Terra (xhigh)
44.7 agentic score
80
fit
| Model | Fit | Agentic | Task cost | Response | Context |
|---|---|---|---|---|---|
| GPT-5.6 Sol (xhigh) OpenAI · Proprietary | 92 | 51.8 | $0.682 | 1.0m | 1.0M |
| GPT-5.6 Sol (max) OpenAI · Proprietary | 89 | 54 | $1.04 | 2.6m | 1.0M |
| GPT-5.6 Sol (high) OpenAI · Proprietary | 88 | 48.5 | $0.453 | 20.2s | 1.0M |
| Claude Fable 5 (Max Effort, Opus 4.8 Fallback) Anthropic · Proprietary | 84 | 52.8 | $2.75 | 1.9m | 1.0M |
| Kimi K3 Kimi · Proprietary | 82 | 50.1 | $0.954 | 1.1m | 1.0M |
| GPT-5.6 Sol (medium) OpenAI · Proprietary | 80 | 44.5 | $0.314 | 18.7s | 1.0M |
| GPT-5.6 Terra (xhigh) OpenAI · Proprietary | 80 | 44.7 | $0.477 | 14.3s | 1.0M |
| Claude Opus 4.8 (max) Anthropic · Proprietary | 79 | 47.2 | $1.80 | 39.3s | 1.0M |
| Grok 4.5 (high) SpaceXAI · Proprietary | 77 | 45.7 | $0.312 | 17.4s | 500K |
| GPT-5.6 Terra (max) OpenAI · Proprietary | 76 | 47.4 | $0.825 | 2.7m | 1.0M |
A small per-token price gap can become a large bill when an agent runs many turns, calls tools, and carries long state. Cost per Intelligence Index task gives a cleaner decision unit than token price alone.
Frontier models are worth it when mistakes are expensive. For routine automation, a cheaper high-fit model can preserve most of the capability while cutting task cost sharply.
Model rankings
Browse the latest ranking pages for overall models, coding, local hardware, open source, Ollama, long context, and agentic workflows.
Live ranking of the best overall AI models by quality, price, speed, and context window.
Current coding leaderboard using LiveCodeBench, Terminal-Bench, and SciCode.
Top open-weight models for self-hosting, Ollama, and low-cost API use.
Best local AI models by hardware tier for self-hosting on Macs, RTX GPUs, and workstations.
Coding-focused local models mapped to 8GB, 24GB, 64GB, and server-class hardware.
Ollama-first picks for coding, chat, reasoning, and low-friction local inference.
Best long-context models for large documents, codebases, and retrieval-heavy workflows.
Rankings for tool use, multi-step execution, and autonomous agent workflows.