๐Ÿ”“Live ranking ยท Editorially updated July 2026

Best Open Source LLM
2026 Ranking + Ollama Guide

The definitive ranking of open-weight AI models you can self-host, fine-tune, and deploy without restrictions. This page now uses the same live dataset and open-source filtering logic as the Explore leaderboard, so the ordering matches the table view across the app.

Self-hostableOllama CompatibleFree to UseFine-tunable

Top 3 Open Source Models

Full Open Source LLM Rankings 2026

RankModelQualityBest PriceTop SpeedMax ContextProviders
1

Muse Spark 1.3

Meta

48.1$2.00/M217 tok/s1M1
2

MiMo-V2.6-Pro

Xiaomi

46.3$0.54/M319 tok/s1M3
3

Qwen3.8 Max

Alibaba

45.4$3.00/M37 tok/s984K1
4

Muse Spark 1.3 (xhigh)

Meta

45.1$2.00/M241 tok/s1M1
5

GLM-5.3

Z AI

44.8$1.70/M420 tok/s1M21
6

Kimi K3

Kimi

43.6$4.80/M317 tok/s1M17
7

GLM 5.3 Flash

Z AI

41.8$0.12/M281 tok/s1M19
8

Qwen3.8 2.4T A95B

Alibaba

39.9$2.85/M178 tok/s1M7
9

Qwen3.8-Flash-Next

Alibaba

39.8$0.23/M54 tok/s1M1
10

DeepSeek V4.1 Flash

DeepSeek

39.5$0.18/M598 tok/s1M21
11

DeepSeek V4 Pro 0813

DeepSeek

36$0.88/M262 tok/s1M11
12

DeepSeek V4 Flash Vision

DeepSeek

34.8$0.32/M228 tok/s1M5

Best Ollama Models 2026

Ollama makes it easy to run open-weight models locally. Here are the top picks by hardware tier:

8GB VRAM

RTX 3070 ยท M2 MacBook Air

  • โ†’ Qwen3 8B (code + chat)
  • โ†’ Gemma 4 12B (multimodal)
  • โ†’ Smaller Qwen3 variants (fast)

16โ€“24GB VRAM

RTX 3090/4090 ยท M2 Pro/Max

  • โ†’ Qwen3-Coder 30B-A3B (coding)
  • โ†’ Qwen3.6 27B (balanced)
  • โ†’ Gemma 4 26B (general)

48GB+ VRAM

2ร— RTX 4090 ยท Mac Studio M2 Ultra

  • โ†’ Qwen3-Coder-Next (local coding)
  • โ†’ Qwen3-Next (efficient MoE)
  • โ†’ Larger current models after size checks

Tip: artifact size is not total runtime memory. Leave room for the context cache, serving framework, operating system, and other applications. The official Qwen3-Coder 30B Ollama artifact is about 19GB; Qwen3-Coder-Next Q4 is about 52GB.

Which Open Source LLM Should You Use?

By Use Case

  • โ†’Coding: Muse Spark 1.3 currently leads the open-source coding subset on this page
  • โ†’Best overall quality: Muse Spark 1.3 is the strongest open-weight pick in the live ranking
  • โ†’Fast responses: DeepSeek V4.1 Flash has the best top-end output speed among the current open-source leaders
  • โ†’Long documents: Muse Spark 1.3 offers the largest context window in this live ranking

By Deployment

  • โ†’Cheap API: GLM 5.3 Flash is currently the cheapest high-ranking option at $0.12/M
  • โ†’Widest availability: GLM-5.3 has the most provider coverage in the current leaderboard
  • โ†’Self-host / Ollama: Use the live ranking here as a shortlist, then jump into the dedicated local and Ollama guides for hardware-specific picks
  • โ†’Commercial use: Verify the exact license for the model family you choose before shipping it into production

How This Ranking Works

Only models with openly available weights are included. This page uses the same grouped Explore dataset and open-source filter as the live leaderboard, then orders those model families by Artificial Analysis Quality Index. We still surface coding benchmarks here when they are available:

MMLU-Pro

Comprehensive knowledge benchmark across 14 domains. Tests breadth of model capability.

AIME 2025

Competition math โ€” tests advanced reasoning. Best signal for math and science tasks.

LiveCodeBench

Contamination-free code generation. Best signal for software development capability.

Compare Open Source Models Side by Side

See live pricing from self-hosting providers, latency, and full benchmark scores for all open source models.

Frequently Asked Questions

What is the best open source LLM in 2026?

Muse Spark 1.3 leads the live open-source ranking with a Quality Index of 48.1. For price-sensitive deployments, GLM 5.3 Flash is the current budget leader at $0.12/M.

What are the best Ollama models in 2026?

For Ollama and other self-hosted setups, start from the strongest live open-weight models here, then narrow by hardware tier. Muse Spark 1.3 is the best option here when context size matters most.

Can open source LLMs match GPT-5 or Claude in 2026?

For most tasks, yes. The top open source models in 2026 trail proprietary leaders by only a small Quality Index margin. The main gaps remain in instruction-following polish, multimodal capability, and very long contexts.

What hardware do I need to run LLMs locally?

As a rule of thumb, 8GB VRAM is enough for smaller 7B to 8B models, 24GB VRAM is a more practical floor for 30B-class models, and 40GB+ is usually required once you move into 70B territory unless you quantize aggressively. Apple Silicon is also viable for smaller and mid-sized open-weight models when unified memory is high enough.

Which open source model is best for coding?

Muse Spark 1.3 is the strongest coding-oriented open source model on this page. For cheap API access, GLM 5.3 Flash is the most economical current pick. See the full coding LLM ranking for more detail.

Related Rankings

Data sources: Rankings use the same live Explore dataset and open-source filter as the methodology page describes. Explore the filtered table in our interactive leaderboard or compare models side by side.