About & data sources
This tool merges pricing and benchmark data for open-source and frontier LLMs into one comparable view. Prices are normalized to USD per 1M tokens (input and output) unless a platform prices differently (GitHub Copilot's current token/AI-Credit rates and legacy request billing are shown on a separate product axis).
Sources
- OpenRouter — model catalog and per-provider endpoint pricing (live API).
- ArtificialAnalysis — Intelligence & Coding indices plus sub-benchmarks (LiveCodeBench, SciCode, Terminal-Bench Hard, τ²-Bench, GPQA, MMLU-Pro) via the v2 API.
- Intelligence.ai / DesignArena — Agentic Web Dev Frontend & Full-Stack Elo leaderboards.
- AWS Bedrock — on-demand token pricing, European regions (eu-central-1 where available).
- Azure AI Foundry — retail token meters plus model-card serving-region checks. A billing/resource region alone is not treated as proof that inference stays in the EU. By company policy, only Azure Direct Global DeepSeek V4 Pro and Kimi K2.7 Code are additionally eligible as EU-hosted equivalents; they remain marked Global because inference may occur outside the EU.
- Google Vertex AI — pay-as-you-go token pricing for Gemini and Model Garden partner models (Claude, Llama, Mistral, DeepSeek, Qwen); offers are marked EU-hosted only where the model supports a documented European serving location.
- TensorX, Inceptron, Scaleway, IONOS, Mistral, Nebius, OVHcloud, STACKIT & T-Systems — direct European serverless/managed catalogs. TensorX is the broadest in-EU host for GLM/Kimi/DeepSeek/MiniMax; Inceptron serves GLM/Kimi/MiniMax from Finland; Scaleway and OVHcloud serve from France; STACKIT and T-Systems publish German/EU catalogs. Nebius is EU-capable but mixes EU, US and UK serving regions, so every model offer is checked separately. See the EU & Sovereign tab.
- GitHub Copilot — current 26-model AI-Credit/token-price catalog plus the 25-model legacy premium-request multiplier table. The UI's separate per-request field applies only to eligible legacy annual Pro/Pro+ plans.
- Anthropic / Claude Code — all 11 currently callable first-party API models, their cache/batch/list prices, active promotions and Claude Code Enterprise terms.
What are “Featured” models?
★ Featured marks the specific models this tool was commissioned to track closely — the current frontier and leading open-weight families that matter most for the price/capability comparison. Filtering to “Featured” (the default on most pages) hides the long tail of older or niche models so the charts and tables stay focused. The featured set is:
- • GPT-5.6 Sol/Terra/Luna, GPT-5.5 & GPT-5.4 (incl. Mini/Nano, low→xhigh effort)
- • Claude Opus 4.8 / 4.7 / 4.6
- • Claude Sonnet 4.6 / 5 & Claude Fable 5
- • Kimi K2.5 / K2.6 / K2.7-Coding
- • GLM 5.1 / 5.2
- • MiniMax M2.5 / M2.7 / M3
- • Xiaomi MiMo-V2.5-Pro
- • DeepSeek V4 Pro
Turn the “Featured” filter off on any page to explore all 803 tracked models. Featured status is derived from the model's family, so every reasoning variant of a featured family (e.g. each GPT-5.5 effort level) is included.
Snapshot
- openrouter: 2026-08-26
- artificialanalysis: 2026-08-26
- designarena: 2026-08-26
- aws_bedrock: 2026-08-26
- azure_foundry: 2026-08-26
- google_vertex: 2026-08-26
- nebius: 2026-08-26
- inceptron: 2026-08-26
- scaleway: 2026-08-26
- ionos: 2026-08-26
- mistral: 2026-08-26
- tensorx: 2026-08-26
- chutes: 2026-08-26
- ovhcloud: 2026-08-26
- stackit: 2026-08-26
- t_systems_llm_hub: 2026-08-26
- aa_coding_agents: 2026-08-26
- github_copilot: 2026-07-22
- claude_code: 2026-07-22
- provider_meta: 2026-07-12
Methodology
The “10:1 blended” cost is (10·input + 1·output) / 11 per 1M tokens, approximating a read-heavy workload. The cheapest such cost across all providers is used for the cost axis. Capability scores are shown as published; AA indices are 0–100, DesignArena values are Elo. The Composite score uses five slots: AA Coding, source-matched AA Coding Agent, AA Intelligence, DesignArena Frontend and DesignArena Full-Stack. AA values are clamped to 0–100. A DesignArena board qualifies at an app-selected minimum of 200 battles, aligned with the source's typical preliminary/reliability threshold; its Elo is converted to the expected score against a fixed Elo 1000 opponent. Each observed slot is converted to its percentile among the current catalog's unique observed values. Every missing slot inherits that model's mean observed percentile, producing a base score exactly equal to the mean of its available percentiles. A final dominance-safe projection prevents missing data from reversing an otherwise unambiguous comparison: when one model covers every reliable slot of another measured model and is no worse in any shared slot, the catalog scores are adjusted by the smallest symmetric amount needed to keep the dominating model at least 0.1 points ahead. The unadjusted base and any adjustment are shown separately in model details. A model with no reliable observed slot receives the neutral fallback 50; its zero evidence coverage remains distinct from a measured score and is excluded from capability charts. Coding Agent results are attached to an exact or explicitly audited model/reasoning identity; every harness result is retained and their median is used. Family-scoped Intelligence.ai / DesignArena results are attached exactly once to the deterministic collapsed-family representative rather than copied to effort siblings. The provenance note explicitly states that this does not identify the tested effort setting. Raw DesignArena score views continue to show Elo. Stable model ids and repositories are preferred over fuzzy names so distinct releases, modes, context tiers and serving routes do not share the wrong price. See the repository README and data/SCRAPING.md for how each source is collected and refreshed.