AI Value Over Time · tracked daily
See what AI actually costs - over time
Every AI benchmark site shows you a snapshot. We track quality-per-dollar and real usage share every day - so you can see whether a model is getting better, cheaper, or more popular, not just where it stands today. No invented index: every number here is a real, sourced quantity or a plain ratio.
- Models tracked
- 291
- Best value (top 10 usage)
- Z.ai: GLM 5.2 (batch)
- History
- 35 days
- Latest data
- August 8, 2026
Quality and usage aren't the whole picture - these track what running a model actually feels like in production, updated every day.
Uptime & latency
Provider reliability
- 1. DeepSeek: DeepSeek V4 Flash 0423100.0%
- 2. Xiaomi: MiMo-V2.598.9%
- 3. Tencent: Hy3100.0%
DeepSeek: DeepSeek V4 Flash 0423 tops the list at 100.0% best-provider uptime.
Full reliability tracker →Real measured throughput
Fastest models
- 1. OpenAI: gpt-oss-120b777 tok/s
- 2. Google: Gemma 4 31B (free)279 tok/s
- 3. MiniMax: MiniMax M2.7215 tok/s
OpenAI: gpt-oss-120b leads at 777 tokens/second via Cerebras.
Full throughput ranking →Daily rank changes
Latest leaderboard movement
- 1. Google: Gemini 3.6 Flash (batch)↓ 11
August 8, 2026: DeepSeek: DeepSeek V4 Flash 0423 led real usage share. Google: Gemini 3.6 Flash (batch) moved 11 places down.
Full daily changelog →Quality per dollar, over time
Top 9 by usage · tracked daily since August 9, 2026
As of August 9, 2026, Z.ai: GLM 5.2 (batch) delivers the best measured quality per dollar among widely-used models, at 489.30 intelligence-index points per dollar ($0.11/1M tokens blended). This chart is tracked daily - first-party history begins August 9, 2026 and grows by one point every night; it is not a static snapshot.
- Z.ai: GLM 5.2 (batch) - 489.30 pts/$ · $0.11/1M
- DeepSeek: DeepSeek V4 Flash 0423 - 460.44 pts/$ · $0.11/1M
- Xiaomi: MiMo-V2.5 - 217.14 pts/$ · $0.18/1M
- MiniMax: MiniMax M3 (batch) - 86.48 pts/$ · $0.52/1M
- DeepSeek: DeepSeek V4 Pro - 83.31 pts/$ · $0.54/1M
- StepFun: Step 3.7 Flash - 70.63 pts/$ · $0.44/1M
- NVIDIA: Nemotron 3 Ultra (free) - 28.37 pts/$ · $1.35/1M
- Anthropic: Claude Opus 4.8 (batch) - 5.73 pts/$ · $10.00/1M
- Anthropic: Claude Opus 4.7 (batch) - 5.50 pts/$ · $10.00/1M
| # | Model | Price ($/1M tokens) | Intelligence index per dollar |
|---|---|---|---|
| 01 | Z.ai: GLM 5.2 (batch) | $0.11/1M | 489.3 |
| 02 | DeepSeek: DeepSeek V4 Flash 0423 | $0.11/1M | 460.4 |
| 03 | Xiaomi: MiMo-V2.5 | $0.18/1M | 217.1 |
| 04 | MiniMax: MiniMax M3 (batch) | $0.52/1M | 86.5 |
| 05 | DeepSeek: DeepSeek V4 Pro | $0.54/1M | 83.3 |
| 06 | StepFun: Step 3.7 Flash | $0.44/1M | 70.6 |
| 07 | NVIDIA: Nemotron 3 Ultra (free) | $1.35/1M | 28.4 |
| 08 | Anthropic: Claude Opus 4.8 (batch) | $10.00/1M | 5.7 |
| 09 | Anthropic: Claude Opus 4.7 (batch) | $10.00/1M | 5.5 |
Daily tracking began Aug 9- this becomes a live chart once there's a third day to plot.
Index date August 8, 2026Data through 00:00 UTC August 9, 2026Calculated 08:38 UTC
How this is calculated
Quality per dollar= Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). No weighting, no composite - a plain ratio of two real, independently sourced numbers, computed fresh from that day's pricing and benchmark snapshot. The Intelligence Index already measures per-response quality (accuracy across a benchmark suite of reasoning, coding, and knowledge tasks) - it is not response volume, so this ratio isn't "cheap model, more tries." A small quality gap next to a huge price gap is what drives a cheap model to the top.
What this ratio can't tell you:
- Consistency - whether a model reliably hits the same quality run to run, not just on average.
- Cost of a wrong answer in your specific use case - a bad response in production code is not the same as a bad response in casual chat.
- Whether a task needs multiple attempts to reach a usable result.
This is a real, industry-standard way to pre-filter candidates by cost-efficiency - not a final verdict. For high-stakes tasks, weigh the absolute Intelligence Index (shown alongside every ratio below) more heavily than the ratio itself, and check a model's measured endpoint uptime for its consistency track record.
Top 15 models by quality per dollar
Intelligence Index points per dollar of blended price - the 15 best-value models among all 85 with a published quality score. This is a ratio, not a capability ranking: an ultra-cheap model with modest absolute quality can outscore a much stronger, pricier one, purely because price sits near zero. Each bar shows its own absolute Intelligence Index and price alongside the ratio, so you can see which is driving the number.
Open vs. closed AI, tracked daily
35 days of real usage history
Open-weight models now handle 74% of real token volume on OpenRouter - up 5 points since July 5, 2026. No other site tracks this as a living, dated trend.
- Open-source
- Closed
Best AI model for…
Ranked by real usage share for that specific task, not overall popularity.
code
Code Generation
DeepSeek: DeepSeek V4 Flash 0423 leads with 20% of task usage, 10 points ahead of OpenAI: GPT-5.6 Luna (batch).
Full code generation ranking, 9 models →general
Content Writing
- 1. DeepSeek: DeepSeek V4 Flash 042322%
- 2. Google: Gemini 2.5 Flash Lite (batch)7%
- 3. Google: Gemini 2.5 Flash (batch)6%
DeepSeek: DeepSeek V4 Flash 0423 leads with 22% of task usage, 16 points ahead of Google: Gemini 2.5 Flash Lite (batch).
Full content writing ranking, 8 models →code
Debugging
DeepSeek: DeepSeek V4 Flash 0423 leads with 19% of task usage, 8 points ahead of Z.ai: GLM 5.2 (batch).
Full debugging ranking, 9 models →general
Translation
- 1. DeepSeek: DeepSeek V4 Flash 042324%
- 2. Google: Gemini 2.5 Flash Lite (batch)17%
- 3. OpenAI: GPT-4o-mini (batch)8%
DeepSeek: DeepSeek V4 Flash 0423 leads with 24% of task usage, 7 points ahead of Google: Gemini 2.5 Flash Lite (batch).
Full translation ranking, 9 models →general
Summarization
- 1. DeepSeek: DeepSeek V4 Flash 042314%
- 2. OpenAI: gpt-oss-120b11%
- 3. Google: Gemini 2.5 Flash (batch)6%
DeepSeek: DeepSeek V4 Flash 0423 leads with 14% of task usage, 4 points ahead of OpenAI: gpt-oss-120b.
Full summarization ranking, 9 models →general
Customer Support
- 1. Google: Gemini 2.5 Flash (batch)20%
- 2. DeepSeek: DeepSeek V4 Flash 042313%
- 3. Google: Gemini 3 Flash Preview (batch)8%
Google: Gemini 2.5 Flash (batch) leads with 20% of task usage, 7 points ahead of DeepSeek: DeepSeek V4 Flash 0423.
Full customer support ranking, 9 models →general
Roleplay & Fiction
DeepSeek: DeepSeek V4 Flash 0423 leads with 47% of task usage, 42 points ahead of DeepSeek: DeepSeek V4 Pro.
Full roleplay & fiction ranking, 9 models →general
Q&A & Knowledge
- 1. DeepSeek: DeepSeek V4 Flash 042320%
- 2. Google: Gemini 3 Flash Preview (batch)5%
- 3. OpenAI: GPT-5.6 Luna (batch)4%
DeepSeek: DeepSeek V4 Flash 0423 leads with 20% of task usage, 16 points ahead of Google: Gemini 3 Flash Preview (batch).
Full q&a & knowledge ranking, 9 models →general
Classification
- 1. DeepSeek: DeepSeek V4 Flash 042315%
- 2. Google: Gemini 2.5 Flash Lite (batch)10%
- 3. OpenAI: GPT-5.6 Luna (batch)7%
DeepSeek: DeepSeek V4 Flash 0423 leads with 15% of task usage, 5 points ahead of Google: Gemini 2.5 Flash Lite (batch).
Full classification ranking, 9 models →Rankings by price and specialty
Comparing a 7B model against a flagship on one scale is meaningless. Each of these is a self-contained "best model" claim within a comparable price tier or use case.
By price
Under $0.50/1M
- 1. Z.ai: GLM 5.2 (batch)Quality 52.6
- 2. OpenAI: GPT-5.6 Luna (batch)Quality 52.3
- 3. DeepSeek: DeepSeek V4 Flash 0423Quality 51.8
$0.50-2/1M
- 1. MiniMax: MiniMax M3 (batch)Quality 45.4
- 2. DeepSeek: DeepSeek V4 ProQuality 45.3
- 3. MoonshotAI: Kimi K2.6Quality 45.1
$2-10/1M
- 1. MoonshotAI: Kimi K3Quality 59.7
- 2. Qwen: Qwen3.8 MaxQuality 58.1
- 3. OpenAI: GPT-5.6 Terra (batch)Quality 56.6
By specialty
Best for coding
- 1. DeepSeek: DeepSeek V4 Flash 042327.4%
- 2. Z.ai: GLM 5.2 (batch)13.7%
- 3. Xiaomi: MiMo-V2.513.2%
Best for agentic work
- 1. DeepSeek: DeepSeek V4 Flash 042338.3%
- 2. Tencent: Hy314.5%
- 3. OpenAI: GPT-5.6 Luna (batch)8.8%
Best value
- 1. Ling-3.0-flash1200.0 pts/$
- 2. inclusionAI: Ling-2.6-flash946.7 pts/$
- 3. Z.ai: GLM 5.2 (batch)489.3 pts/$
Daily snapshots
Every published day gets its own permanent, unchanging permalink - a verifiable point in time for citations, not a page that silently drifts.
How this is calculated
No composite index. No invented weights. Two plain, independently-checkable numbers, each plotted as a time series instead of a single snapshot:
Quality per dollar
Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). A plain division, recomputed from that day's real pricing and benchmark snapshot.
Usage share
Real token volume through OpenRouter, as a share of that day's total - no sampling, no estimation.
We tried composite indices three times this year (a weighted-sum composite, a weighted-geometric-mean composite, and a usage-vs-quality divergence index) and killed all three after finding real defensibility problems - see the full methodology page for that history. The lesson: publish real numbers over time, not invented ones.
Source: OpenRouter (openrouter.ai/rankings), as of August 9, 2026.
Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).
Methodology: smophy.ai/benchmark/methodology
Token counts originate from each provider's own tokenizer and are not directly comparable across providers.
Journalist or researcher? Pre-computed citable stats and CSV/JSON exports →
