Live data
Assembling the frontier
Ranking the latest models and provider endpoints.
Loading model dataLive data
Ranking the latest models and provider endpoints.
Loading model dataSide by side
Select up to 4 models to analyze quality, performance, pricing, and benchmarks.
Quick picks from top 10 models by Quality Index
Highest Quality
GPT-5.6 Luna (xhigh)
34.8
Fastest Output
GPT-5.6 Luna (xhigh)
116 tok/s
Best Value
GPT-5.6 Luna (xhigh)
$0.45/M
Detailed benchmark scores across evaluations
Scores match the exact selected model variant. Current Artificial Analysis values take priority; curated references fill missing metrics and are labeled below. Missing values are shown as —.
| Benchmark | GPT-5.6 Luna (xhigh) |
|---|---|
GDPval-AA Elo🤖 Professional-work task Elo rating Source: Artificial Analysis ↗ | 1428👑Artificial Analysis |
GDPval-AA🤖 Professional-work task benchmark Source: Artificial Analysis ↗ | 46%👑Artificial Analysis |
Long Context Recall📄 Artificial Analysis long-context reasoning score Source: Artificial Analysis ↗ | 82%👑Artificial Analysis |
Humanity's Last Exam🧠 Hard, anti-saturation reasoning exam Source: Scale AI ↗ | 37%👑Artificial Analysis |
GPQA Diamond🧠 PhD-level science reasoning Source: Google ↗ | 90%👑Artificial Analysis |
SciCode💻 Scientific coding benchmark Source: SciCode ↗ | 51%👑Artificial Analysis |
CritPt🧠 Critical point / robustness benchmark Source: CritPt ↗ | 21%👑Artificial Analysis |
AA-Omniscience Accuracy | 42%👑Artificial Analysis |
| Metric | GPT-5.6 Luna (xhigh) |
|---|---|
| Creator | OpenAI |
| Quality Index | 34.8👑 |
| Price per Million Tokens | $0.45💰 |
| Output Speed | 116 tok/s⚡ |
| Context Window | 1.0M📚 |
| Latency (TTFT) | 31.59s🚀 |
| Provider | OpenAI1 more providers |
Estimated cost based on 1 million tokens per day usage
GPT-5.6 Luna (xhigh)
$14
/month
Use coding, reasoning, math, and tool-use benchmarks to see where a model is actually strong instead of relying on a single overall score. A model that leads in quality may still be wrong for your workflow if your primary constraint is latency or cost.
The same model can be cheap on one host and expensive on another, or fast on one provider and unusable on the next. If the model looks promising, move to provider comparison before you commit.
Ranking library
Focused rankings for the decisions engineers actually make.