Are Kimi K3 weights available?
Yes. Moonshot AI has released the Kimi K3 weights on Hugging Face. The model remains impractical for ordinary local hardware because of its 2.8T-parameter sparse architecture.
Live data
Ranking the latest models and provider endpoints.
Loading model dataKimi · Open weights
Configuration: Kimi K3 (max)
Endpoint: Makora. The headline price, speed, latency, and context use this one endpoint. It has the lowest known 3:1 blended price for this configuration.
Family rank uses the strongest observed Intelligence Index configuration in the benchmark dataset. It stays the same when you select a different effort setting.
Benchmark data: Using Artificial Analysis API benchmark data, cached for up to 24 hours. The source measurement date is not supplied.
Provider data: Provider data fetched 7 Sept 2026, 13:50 UTC. Benchmark responses may be cached for up to 24 hours. How to read these dates.
The answer in 60 seconds
Kimi K3 now pairs frontier-band coding with released weights and a one-million-token window. That is a major openness milestone, but not a laptop story: its 2.8T sparse architecture still requires datacenter-scale storage, memory, and interconnect. Hosted K3 is the practical choice for most teams; self-hosting is for organizations that need control badly enough to fund the infrastructure.
Independent data
Kimi K3 (max) Benchmark scores are for this exact source label. A dash means the metric is unavailable; it does not mean zero.
Price, speed, and time to first token below come from Makora.
Kimi K3 (max). Different tests and benchmark versions measure different jobs; compare the same version and configuration.
Choose an effort setting to inspect its benchmarks and providers. Each row's price and speed come from its named endpoint, selected by lowest known blended price.
| Configuration | Intelligence | Coding | Agentic | Endpoint | Blended / 1M | Output speed |
|---|---|---|---|---|---|---|
| Kimi K3 (max)Selected | 50.2 | 76.2 | 50.9 | Makora | $5.10 | 26.0 tok/s |
| Kimi K3 (low) | 38.8 | 72.0 | — | Baseten | $6.00 | 101.2 tok/s |
Where to run it
Kimi K3 (max). Each row describes one observed endpoint. Prices are USD per million tokens; the blended estimate uses three input tokens per output token. Availability and serving limits can differ by provider.
| Provider / endpoint | Input / 1M | Output / 1M | Blended / 1M | Output speed | TTFT | Context |
|---|---|---|---|---|---|---|
| MakoraShown above | $2.55 | $12.75 | $5.10 | 26.0 tok/s | 2.664s | 1.05M |
| Bitdeer AI | $2.66 | $13.30 | $5.32 | 87.0 tok/s | 1.483s | 262K |
| DigitalOcean | $2.85 | $14.25 | $5.70 | 42.7 tok/s | 0.718s | 1.05M |
| Baseten | $3.00 | $15.00 | $6.00 | 89.4 tok/s | 0.666s | 1.05M |
| Databricks | $3.00 | $15.00 | $6.00 | 177.6 tok/s | 0.732s | 205K |
| Fireworks | $3.00 | $15.00 | $6.00 | 114.3 tok/s | 0.737s | 1.05M |
| Kimi | $3.00 | $15.00 | $6.00 | 42.2 tok/s | 2.781s | 1.05M |
| Modal | $3.00 | $15.00 | $6.00 | 149.4 tok/s | 0.469s | 1.05M |
| Nebius | $3.00 | $15.00 | $6.00 | 95.2 tok/s | 1.140s | 1.05M |
| Parasail | $3.00 | $15.00 | $6.00 | 95.2 tok/s | 0.688s | 1.05M |
| Together AI | $3.00 | $15.00 | $6.00 | 64.2 tok/s | 0.937s | 1.05M |
| Fireworks (FAST) | $4.50 | $22.50 | $9.00 | 111.5 tok/s | 0.737s | 1.05M |
| Inco (FAST) | $6.00 | $30.00 | $12.00 | 259.5 tok/s | 0.561s | 1.05M |
Provider data fetched 7 Sept 2026, 13:50 UTC. Benchmark responses may be cached for up to 24 hours. Missing prices, speed, latency, or context are shown as a dash. Explore all endpoints.
Source truth
Editorial analysis
We have not claimed hands-on use here. The analysis below combines provider-confirmed facts with the independent data shown above; each section labels that distinction.
Boundaries
Alternatives
Near-Max Qwen capability with downloadable weights—and a 2.4T-scale infrastructure bill hidden behind the word “open.”
Configuration: Qwen3.8 2.4T A95B
02The top-10 outlier: open weights, native vision, near-frontier agent scores, and Flash-class serving economics in the same release.
Configuration: GLM-5.3-Flash
03The broad API flagship: frontier coding and tool use with more controllable effort than its single leaderboard row suggests.
Configuration: GPT-5.6 Sol (max)
Evidence register
Checked August 28, 2026. Pricing and availability should be verified again before procurement.
Direct answers
Yes. Moonshot AI has released the Kimi K3 weights on Hugging Face. The model remains impractical for ordinary local hardware because of its 2.8T-parameter sparse architecture.
Not in the usual desktop sense. Kimi recommends supernode configurations with at least 64 accelerators, and the weight files alone are measured at datacenter scale.