The interesting number is agentic value, not the tie on overall score
A leaderboard tie can hide the commercial difference. Grok 4.6 combines a frontier-band Agentic Index with a lower blended token price than the premium Anthropic and OpenAI tiers. If an agent makes many calls, retries, and tool decisions, that price gap compounds faster than a small movement in the overall index.
The right test is a completed-task harness with tool costs and retries included. A cheap token that causes another loop is not cheap; a lower-priced model that finishes in fewer steps is.