Meta is back in the AI race
Muse Glimmer puts Apache 2.0 weights back inside Metaβs comeback. Spark is chasing the frontier; Glimmer is bringing agents home.
Read articleBenchmark analysis, model rankings, and practical guides for builders.
Muse Glimmer puts Apache 2.0 weights back inside Metaβs comeback. Spark is chasing the frontier; Glimmer is bringing agents home.
Read articleThe definitive guide to Moonshot AI's 2.8T model: verified specifications, benchmark caveats, $3/$15 API pricing, 1M context, architecture, access, code examples, and who should use it.
Open weights are roughly four months behind the aggregate frontier. Across coding, agents, reasoning, cost, privacy, and deployment, that average hides the decision that actually matters.
SWE-bench, Terminal-Bench, LiveCodeBench, and SciCode rank different jobs. A practical guide to choosing the benchmark that matches the code you need written.
Price per token ignores retries. This transparent cost-per-solved-task framework adds terminal success, expected attempts, latency, review, and escalation risk.
Artificial Analysis Intelligence Index v4.1 changes the model-selection question from "which model is smartest?" to "what does a completed agent task cost?" A data-backed guide to Agentic Index, cost per task, cache pricing, Terminal-Bench 2.1, ΟΒ³-Bench Banking, and GDPval-AA v2.
After April broke the ceiling (GPT-5.5 at 60.24, Opus 4.7, DeepSeek V4, Kimi K2.6), May went quiet on scale and loud on architecture. SubQ shipped the first commercial subquadratic LLM with a 12M context. Zyphra dropped an 8B MoE trained on AMD. OpenAI made GPT-5.5 Instant the new ChatGPT default.
A long letter on why I built AI Bottlenecks. EUV machines, indium phosphide wafers, CoWoS slots, gas turbines, and the four physical chokepoints gating the 2026 to 2030 AI buildout. The constraint is never at the loud end of the stack.
A synthesis of Artificial Analysis's 2025 Year-End State of AI report. Reasoning models took the leaderboard, the cost of intelligence collapsed 100x, coding agents went mainstream, and the frontier got more contested β not less.
DeepSeek V4-Pro and V4-Flash just dropped. 1.6T MoE, native 1M context, MIT weights, trained on Huawei Ascend. Priced at ~1/20th of Opus 4.7. The model Jensen warned about eight days ago on Dwarkesh β live on Hugging Face.
Moonshot AI just shipped Kimi K2.6. 1T parameters, 262K context, 4,000 tool calls in a single run, and benchmarks that put it shoulder to shoulder with GPT-5.4 and Claude Opus 4.6.
Claude Mythos is locked behind a 50-company firewall. GLM-5.1 beat GPT-5.4 on coding under MIT license. Gemma 4 went Apache 2.0. The full breakdown of AI's first week of April.
GPT-5.4 matched Gemini 3.1 Pro within 0.01 points. NVIDIA unveiled trillion-parameter infrastructure. Anthropic clashed with the Pentagon. Nine text models shipped, seven open-weight. Full breakdown.
Google pushed a mid-cycle update to Gemini 3 Pro on February 11. We measured every benchmark delta: +4.2pp on SWE-Bench, +5.2pp on AIME 2025, #1 on LM Arena. Here is what changed and who should switch.
AI is not destroying jobs. It is exposing that most white-collar work was never truly meaningful. The K-shaped economy is splitting fast, agents are hiring humans, and the next 2-3 years decide which side you land on.
A data-driven DeepSeek hub that connects benchmarks to provider reality: pricing, speed, and time to first token across hosts.
A MiniMax hub for builders. Understand M1 vs M2 vs M2.1 and compare providers by price, speed, and time to first token.
The narrative around AI has long revolved around compute power. As we enter 2026, a quieter shift is underway. Memory, not chips, is becoming the constraint.
Twelve months ago, we were debating whether AI could reason. Now we're debating who owns the reasoning.
DeepSeek V3.2 hit 96% on AIME 2025. Xiaomi dropped a frontier model. GLM-4.7 claimed the coding crown.
OpenAI declared "Code Red" in December 2025. China's 15 open-weight models, the efficiency revolution, and custom silicon converged.
Analysis of 114 models reveals a market in transition: benchmarks saturating, open-weight matching proprietary at 10x lower cost.
The AI arms race explodes in November 2025 with three frontier releases in 12 days. Benchmarks, pricing, and where each dominates.
Moonshot AI's open-weight MoE takes on OpenAI's proprietary GPT-5.1. Architecture, benchmarks, pricing.
Moonshot AI's K2 Thinking lands a 67 on Artificial Analysis, sets agentic records, proves open weights compete on reasoning.
DeepSeek V3.1, Qwen3-235B, GLM-4.6 - performance benchmarks, pricing, and deployment insights.
Lessons from Andrej Karpathy: model collapse, low-entropy outputs, and how to push past the slop.
Z.ai's GLM-4.5 hybrid reasoning vs Moonshot AI's Kimi-K2 1T parameter architecture.
Z.ai's GLM-4.5 vs Alibaba's Qwen3-235B with massive parameter count and FP8 optimization.
Compare 100+ LLMs by price, speed, and benchmarks.