🤖Live agentic ranking · AA Intelligence Index v4.1 · Updated June 2026

Best Agentic AI Models
2026 ranking for tool use and autonomous work

Rankings here prioritize Agentic Index, terminal tasks, tool-use benchmarks, and cost per Intelligence Index task where available, so the order reflects real workflow tradeoffs instead of generic chat quality alone.

What changed in the agentic index

Harder agent tasks

The ranking now leans harder into Terminal-Bench, τ-Bench, and GDPval-AA style workflows that test longer agent trajectories.

Cost per task

Token price alone misses long-horizon agent cost. Cost per Intelligence Index task gives a better unit-economics view.

Cache-aware pricing

Cached input pricing now matters for agent loops with repeated context, especially when prompts and tool traces get large.

Agentic fit lab

Find the right agent model

Task

Autonomy needed

Context load

Optimize for

Frontier map

Agentic Index vs task cost

14 high-fit models

112 with cost/task

0112233435465$0.010$0.030$0.100$0.300$1.00$3.00Agentic IndexCost per taskGPT-5.6 Sol (xhigh) · Agentic 51.8 · $0.682 per taskGPT-5.6 Sol (max) · Agentic 54 · $1.04 per taskGPT-5.6 Sol (high) · Agentic 48.5 · $0.453 per taskClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) · Agentic 52.8 · $2.75 per taskKimi K3 · Agentic 50.1 · $0.954 per taskGPT-5.6 Sol (medium) · Agentic 44.5 · $0.314 per taskGPT-5.6 Terra (xhigh) · Agentic 44.7 · $0.477 per taskClaude Opus 4.8 (Adaptive Reasoning, Max Effort) · Agentic 47.2 · $1.80 per taskGrok 4.5 (high) · Agentic 45.7 · $0.312 per taskGPT-5.6 Terra (max) · Agentic 47.4 · $0.825 per taskGPT-5.5 (xhigh) · Agentic 44.9 · $0.993 per taskGPT-5.5 (high) · Agentic 43.5 · $0.668 per taskGLM-5.2 (max) · Agentic 43.1 · $0.468 per taskGPT-5.6 Luna (max) · Agentic 45.6 · $0.209 per taskClaude Opus 4.7 (Adaptive Reasoning, Max Effort) · Agentic 44.4 · $1.97 per taskGPT-5.6 Terra (high) · Agentic 41.3 · $0.336 per taskGPT-5.6 Sol (low) · Agentic 40 · $0.198 per taskGPT-5.6 Luna (xhigh) · Agentic 42.9 · $0.139 per taskClaude Sonnet 5 (Adaptive Reasoning, Max Effort) · Agentic 46.7 · $1.53 per taskGPT-5.5 (medium) · Agentic 37.8 · $0.406 per taskGPT-5.6 Luna (high) · Agentic 40.1 · $0.095 per taskGemini 3.5 Flash (high) · Agentic 37.4 · $0.586 per taskMuse Spark 1.1 (xhigh) · Agentic 37.5 · $0.261 per taskClaude Sonnet 4.6 (Adaptive Reasoning, Max Effort) · Agentic 40.8 · $1.14 per taskGPT-5.6 Terra (medium) · Agentic 37 · $0.175 per taskDeepSeek V4 Pro (Reasoning, Max Effort) · Agentic 36.4 · $0.045 per taskMiniMax-M3 · Agentic 35.4 · $0.125 per taskGPT-5.6 Sol (Non-reasoning) · Agentic 34.9 · $0.200 per taskDeepSeek V4 Pro (Reasoning, High Effort) · Agentic 34.4 · $0.041 per taskClaude Sonnet 5 (Non-reasoning, High Effort) · Agentic 33.7 · $0.374 per taskQwen3.7 Max · Agentic 30.6 · $1.03 per taskGPT-5.5 (low) · Agentic 30.4 · $0.211 per taskGPT-5.6 Terra (low) · Agentic 30.6 · $0.154 per taskGPT-5.4 mini (xhigh) · Agentic 30.2 · $0.452 per taskMiMo-V2.5-Pro · Agentic 29.1 · $0.031 per taskDeepSeek V4 Flash (Reasoning, Max Effort) · Agentic 31.1 · $0.022 per taskKimi K2.6 · Agentic 30.3 · $0.335 per taskGLM-5.1 (Reasoning) · Agentic 29.9 · $0.226 per taskGPT-5.6 Luna (medium) · Agentic 31 · $0.050 per taskGPT-5.4 nano (xhigh) · Agentic 27.5 · $0.133 per taskGPT-5.6 Terra (Non-reasoning) · Agentic 29.3 · $0.179 per taskGemini 3.1 Pro Preview · Agentic 21.4 · $0.291 per taskGPT-5.5 (Non-reasoning) · Agentic 25.8 · $0.171 per taskGrok Build 0.1 0616 · Agentic 28 · $0.212 per taskNemotron 3 Ultra 550B A55B (Reasoning) · Agentic 27.4 · $0.244 per taskQwen3.6 Plus · Agentic 27.6 · $0.314 per taskMiMo-V2.5 · Agentic 23.7 · $0.010 per taskQwen3.6 27B (Reasoning) · Agentic 27 · $0.265 per taskClaude 4.5 Sonnet (Reasoning) · Agentic 24.6 · $0.413 per taskGPT-5.6 Luna (low) · Agentic 25.4 · $0.041 per taskQwen3.7 Plus · Agentic 20.8 · $0.207 per taskGLM-4.7 (Reasoning) · Agentic 25.4 · $0.323 per taskGrok 4.3 (high) · Agentic 24.1 · $0.139 per taskGPT-5.1 (high) · Agentic 21 · $0.270 per taskQwen3.6 27B (Non-reasoning) · Agentic 23.3 · $0.359 per taskGPT-5 (high) · Agentic 25.7 · $0.238 per taskQwen3.5 122B A10B (Reasoning) · Agentic 20.7 · $0.241 per taskQwen3.5 397B A17B (Reasoning) · Agentic 19.8 · $0.333 per taskQwen3.6 35B A3B (Reasoning) · Agentic 21.4 · $0.179 per taskGPT-5.6 Luna (Non-reasoning) · Agentic 22 · $0.055 per taskMistral Medium 3.5 · Agentic 19 · $0.563 per taskRing-2.6-1T · Agentic 18.9 · $0.345 per taskGrok 4.3 (Non-reasoning) · Agentic 22.8 · $0.293 per taskQwen3.5 122B A10B (Non-reasoning) · Agentic 15.8 · $0.177 per taskGLM-4.6 (Reasoning) · Agentic 17.7 · $0.283 per taskClaude 4.5 Haiku (Reasoning) · Agentic 16.4 · $0.237 per taskMercury 2 · Agentic 9.6 · $0.075 per taskMistral Medium 3.1 · Agentic 6.2 · $0.145 per taskNova 2.0 Pro Preview (medium) · Agentic 7 · $0.173 per taskDeepSeek R1 (Jan '25) · Agentic 3.1 · $0.247 per taskMistral Small 3.1 · Agentic 5.2 · $0.041 per taskGemini 2.5 Pro · Agentic 7.1 · $0.198 per taskQwen3.5 9B (Reasoning) · Agentic 7.4 · $0.164 per taskHyperNova 60B 2605 · Agentic 6.7 · $0.018 per taskgpt-oss-20b (high) · Agentic 3.1 · $0.018 per taskMistral Small 4 (Reasoning) · Agentic 4.7 · $0.098 per taskMistral Small 3.2 · Agentic 2 · $0.124 per taskDeepSeek V3 (Dec '24) · Agentic 1.6 · $0.023 per taskMistral Large 3 · Agentic 5.5 · $0.060 per taskGemini 3.1 Flash-Lite · Agentic 6.2 · $0.042 per taskGPT-5.6 Sol (xhigh)GPT-5.6 Sol (max)GPT-5.6 Sol (high)

GPT-5.6 Sol (xhigh)

OpenAI

Agentic

51.8

Task cost

$0.682

Top 3 agentic models

Full agentic model ranking

The models below are ranked for autonomous execution, tool use, multi-step reliability, and task-level economics.

RankModelAgenticCost / taskTerminal-Benchτ-BenchGDPval-AAResponse
1

GPT-5.6 Sol (max)

OpenAI

54$1.0466%85%17432.6m
2

Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

Anthropic

52.8$2.7563%99%17601.9m
3

Kimi K3

Kimi

50.1$0.954N/AN/A16841.1m
4

GPT-5.6 Terra (max)

OpenAI

47.4$0.82558%86%15812.7m
5

Claude Opus 4.8 (Adaptive Reasoning, Max Effort)

Anthropic

47.2$1.8058%94%160039.3s
6

Claude Sonnet 5 (Adaptive Reasoning, Max Effort)

Anthropic

46.7$1.53N/AN/A16072.7m
7

Grok 4.5 (high)

SpaceXAI

45.7$0.312N/AN/A153517.4s
8

GPT-5.6 Luna (max)

OpenAI

45.6$0.209N/AN/A15841.4m
9

Gemini 3.5 Flash (medium)

Google

45.4N/A39%96%N/A19.3s
10

GPT-5.5 (xhigh)

OpenAI

44.9$0.99361%94%14931.4m

How to choose a model for agents

If your agents browse, call APIs, run tools, or plan several steps ahead, start here rather than on a generic leaderboard. The best agentic model is the one that stays reliable under execution at a cost and response-time profile you can actually ship.

Once you have a shortlist, open the finalists on Compare and validate whether the provider you want can deliver the right price and latency profile.

Related live rankings

Agentic performance overlaps with coding and long-context performance, but it is not the same thing. Use the related pages below if your use case is specialized around software work, model routing, or large-document reasoning.

Frequently Asked Questions

What is the best AI model for agents in 2026?

The live top-ranked model on this page is the best starting point. The right answer depends on whether you prioritize raw capability, reliability under tool use, or latency in production. Check the ranking table above for the current leader.

What makes an LLM agentic?

Agentic models can maintain plans across many steps, call tools reliably, follow complex instructions, and recover gracefully when a workflow hits an unexpected state.

Are coding models automatically good for agents?

Not always. Coding strength helps with tool-writing and structured output, but agentic performance also requires strong planning, tool orchestration, and error recovery.

How should I validate an agentic model shortlist?

Use benchmarks to get a shortlist, then run the finalists on the actual tasks your agent will perform. Production testing on real workflows is the only way to make the final call.

Does latency matter for agentic AI?

Yes — significantly. In multi-step agents, slow thinking compounds across many tool calls. A model that is 30% slower can double the wall-clock time of a complex workflow.

What benchmarks predict agentic performance?

Terminal-Bench, τ-Bench, and GDPval-AA are strong predictors of real agentic capability. IFBench remains useful as a supplemental instruction-following signal, but it is no longer the main filter for frontier models.