Live data
Assembling the frontier
Ranking the latest models and provider endpoints.
Loading model dataLive data
Ranking the latest models and provider endpoints.
Loading model dataSide by side
Select up to 4 models to analyze quality, performance, pricing, and benchmarks.
Quick picks from top 10 models by Quality Index
Highest Quality
Muse Spark 1.3 (max)
48.2
Fastest Output
Muse Spark 1.3 (max)
274 tok/s
Best Value
Muse Spark 1.3 (max)
$2.00/M
Detailed benchmark scores across evaluations
Scores match the exact selected model variant. Current Artificial Analysis values take priority; curated references fill missing metrics and are labeled below. Missing values are shown as —.
| Benchmark | Muse Spark 1.3 (max) |
|---|---|
GDPval-AA Elo🤖 Professional-work task Elo rating Source: Artificial Analysis ↗ | 1703👑Artificial Analysis |
GDPval-AA🤖 Professional-work task benchmark Source: Artificial Analysis ↗ | 60%👑Artificial Analysis |
Long Context Recall📄 Artificial Analysis long-context reasoning score Source: Artificial Analysis ↗ | 83%👑Artificial Analysis |
Humanity's Last Exam🧠 Hard, anti-saturation reasoning exam Source: Scale AI ↗ | 49%👑Artificial Analysis |
GPQA Diamond🧠 PhD-level science reasoning Source: Google ↗ | 94%👑Artificial Analysis |
SciCode💻 Scientific coding benchmark Source: SciCode ↗ | 59%👑Artificial Analysis |
CritPt🧠 Critical point / robustness benchmark Source: CritPt ↗ | 25%👑Artificial Analysis |
AA-Omniscience Accuracy | 44%👑Artificial Analysis |
| Metric | Muse Spark 1.3 (max) |
|---|---|
| Creator | Meta |
| Quality Index | 48.2👑 |
| Price per Million Tokens | $2.00💰 |
| Output Speed | 274 tok/s⚡ |
| Context Window | 1.0M📚 |
| Latency (TTFT) | 24.21s🚀 |
| Provider | Meta |
Estimated cost based on 1 million tokens per day usage
Muse Spark 1.3 (max)
$60
/month
Use coding, reasoning, math, and tool-use benchmarks to see where a model is actually strong instead of relying on a single overall score. A model that leads in quality may still be wrong for your workflow if your primary constraint is latency or cost.
The same model can be cheap on one host and expensive on another, or fast on one provider and unusable on the next. If the model looks promising, move to provider comparison before you commit.
Ranking library
Focused rankings for the decisions engineers actually make.