๐Ÿ“Š Methodology & Data Sources

How WhatLLM.org Compares LLMs

Transparent documentation of our data sources, the Artificial Analysis Intelligence Index, and how WhatLLM.org helps you find the right LLM for your needs.

Primary Data Source: Artificial Analysis

WhatLLM.org's benchmark data and Intelligence Index scores come from Artificial Analysis (artificialanalysis.ai), an independent research organization that rigorously benchmarks LLMs across quality, speed, and price.

We fetch data from Artificial Analysis and retain a dated snapshot when API data is unavailable. Visit their site for the complete benchmark methodology and latest research.

Data freshness and editorial review

Provider data is fetched on demand and reused for up to one hour in each server process. Benchmark responses may be cached for up to 24 hours. A daily refresh job is scheduled. If a fetch fails, WhatLLM uses a dated saved snapshot.

A fetch date records when WhatLLM obtained a provider payload. A snapshot date records when the saved fallback was exported. Neither tells you when Artificial Analysis ran each benchmark. Where a cached benchmark response does not expose its fetch time, we leave that date unknown. Refresh intervals describe our cache policy; they do not guarantee a new measurement on every refresh.

Editorial reviews are separate from data retrieval. This explanation was reviewed on . Model guides and decision pages show their own review dates. Check the original provider documentation before relying on a price or availability claim.

WhatLLM editorial content is available under CC BY 4.0. Third-party benchmark, price, and availability data remains subject to Artificial Analysis and provider terms. Our public model data includes source status, available timestamps, attribution, and field definitions.

The Artificial Analysis Intelligence Index

The Artificial Analysis Intelligence Index combines several evaluations into an overall capability score. Its benchmark set and weights change between versions. WhatLLM displays the supplied score and version where available; compare results from the same version and use task-specific evidence for your workload.

โš ๏ธ Attribution: WhatLLM.org does not calculate or create the Intelligence Index. We display Artificial Analysis's scores to help users compare models. For the complete mathematical methodology, visit Artificial Analysis's intelligence benchmarking methodology.

Read the score with its methodology version

The official methodology reviewed on 6 September 2026 describes a weighted average of agentic, coding, general, and scientific-reasoning evaluations. The source methodology and version history specify the evaluation set, weights, and testing procedures for each release.

A saved WhatLLM snapshot can contain an earlier index version than the latest source website. A changed score can reflect a methodology change as well as a model change. Check the source version before comparing scores across dates; a missing version remains unknown.

Examples of benchmarks in model comparisons

These examples explain scores you may encounter in current or historical model comparisons. They are not a list of the current Intelligence Index components. Availability, benchmark version, and evaluation settings vary; consult the linked source methodology for current inclusion.

GPQA Diamond

Reasoning

PhD-level science questions requiring multi-step reasoning. Tests deep domain expertise in physics, chemistry, and biology.

AIME 2025

Mathematics

American Invitational Mathematics Examination problems. Tests advanced mathematical problem-solving.

LiveCodeBench

Coding

Real-world coding challenges with execution-based evaluation. Tests practical programming ability.

SWE-Bench Verified

Software Engineering

Real GitHub issues requiring code changes. Tests ability to understand codebases and implement fixes.

MMLU-Pro

Knowledge

A challenging multitask language-understanding evaluation covering academic and professional knowledge.

Humanity's Last Exam

Frontier Reasoning

Expert-crafted questions designed to push the limits of AI reasoning.

ฯ„ยฒ-Bench Telecom

Agentic

Complex multi-step agent tasks. Tests tool use and planning capabilities.

What WhatLLM.org Adds

While Artificial Analysis provides the underlying data and Intelligence Index, WhatLLM.org offers additional value through:

Additional Data Sources

Beyond Artificial Analysis, WhatLLM.org cross-references data from:

Update Frequency

New model coverage depends on availability in the source data and editorial review. We do not guarantee same-day coverage of a release. The data freshness section explains the provider, benchmark, and saved-snapshot refresh policies.

Limitations & Transparency

We believe in honest representation of our data:

How to Cite

When citing WhatLLM.org:

Bristot, D. (2025). WhatLLM.org. https://whatllm.org

When citing the Intelligence Index methodology:

Artificial Analysis. (2025). Artificial Analysis Intelligence Index. https://artificialanalysis.ai

For academic work using Intelligence Index scores, we recommend citing Artificial Analysis as the primary source for the methodology and data.

Frequently Asked Questions

Where does WhatLLM.org get its data?

Primarily from Artificial Analysis, supplemented by official model papers and benchmark leaderboards. The Intelligence Index scores are calculated by AA using their published methodology.

What is the Intelligence Index?

The Artificial Analysis Intelligence Index combines several evaluations into an overall capability score. Its benchmark set and weights change between versions. WhatLLM displays the supplied score and version where available; compare results from the same version and use task-specific evidence for your workload.

What value does WhatLLM.org add?

Interactive visualization, advanced filtering, side-by-side comparisons, in-depth blog analysis, and a simplified interface for quickly identifying the best LLM for your use case.

How often is the data updated?

Provider data is fetched on demand and reused for up to one hour in each server process. Benchmark responses may be cached for up to 24 hours. A daily refresh job is scheduled. If a fetch fails, WhatLLM uses a dated saved snapshot. Fetch dates and editorial review dates are separate; neither is a benchmark measurement date.

Next Steps