SmophyAI

Published daily · Next update 02:00 UTC

AI Value Over Time · tracked daily

See what AI actually costs - over time

Every AI benchmark site shows you a snapshot. We track quality-per-dollar and real usage share every day - so you can see whether a model is getting better, cheaper, or more popular, not just where it stands today. No invented index: every number here is a real, sourced quantity or a plain ratio.

Models tracked
291
Best value (top 10 usage)
Z.ai: GLM 5.2 (batch)
History
35 days
Latest data
August 8, 2026

Quality and usage aren't the whole picture - these track what running a model actually feels like in production, updated every day.

Uptime & latency

Provider reliability

  1. 1. DeepSeek: DeepSeek V4 Flash 0423100.0%
  2. 2. Xiaomi: MiMo-V2.598.9%
  3. 3. Tencent: Hy3100.0%

DeepSeek: DeepSeek V4 Flash 0423 tops the list at 100.0% best-provider uptime.

Full reliability tracker

Real measured throughput

Fastest models

  1. 1. OpenAI: gpt-oss-120b777 tok/s
  2. 2. Google: Gemma 4 31B (free)279 tok/s
  3. 3. MiniMax: MiniMax M2.7215 tok/s

OpenAI: gpt-oss-120b leads at 777 tokens/second via Cerebras.

Full throughput ranking

Daily rank changes

Latest leaderboard movement

  1. 1. Google: Gemini 3.6 Flash (batch)↓ 11

August 8, 2026: DeepSeek: DeepSeek V4 Flash 0423 led real usage share. Google: Gemini 3.6 Flash (batch) moved 11 places down.

Full daily changelog

Quality per dollar, over time

Top 9 by usage · tracked daily since August 9, 2026

As of August 9, 2026, Z.ai: GLM 5.2 (batch) delivers the best measured quality per dollar among widely-used models, at 489.30 intelligence-index points per dollar ($0.11/1M tokens blended). This chart is tracked daily - first-party history begins August 9, 2026 and grows by one point every night; it is not a static snapshot.

  1. Z.ai: GLM 5.2 (batch) - 489.30 pts/$ · $0.11/1M
  2. DeepSeek: DeepSeek V4 Flash 0423 - 460.44 pts/$ · $0.11/1M
  3. Xiaomi: MiMo-V2.5 - 217.14 pts/$ · $0.18/1M
  4. MiniMax: MiniMax M3 (batch) - 86.48 pts/$ · $0.52/1M
  5. DeepSeek: DeepSeek V4 Pro - 83.31 pts/$ · $0.54/1M
  6. StepFun: Step 3.7 Flash - 70.63 pts/$ · $0.44/1M
  7. NVIDIA: Nemotron 3 Ultra (free) - 28.37 pts/$ · $1.35/1M
  8. Anthropic: Claude Opus 4.8 (batch) - 5.73 pts/$ · $10.00/1M
  9. Anthropic: Claude Opus 4.7 (batch) - 5.50 pts/$ · $10.00/1M
#ModelPrice ($/1M tokens)Intelligence index per dollar
01Z.ai: GLM 5.2 (batch)$0.11/1M489.3
02DeepSeek: DeepSeek V4 Flash 0423$0.11/1M460.4
03Xiaomi: MiMo-V2.5$0.18/1M217.1
04MiniMax: MiniMax M3 (batch)$0.52/1M86.5
05DeepSeek: DeepSeek V4 Pro$0.54/1M83.3
06StepFun: Step 3.7 Flash$0.44/1M70.6
07NVIDIA: Nemotron 3 Ultra (free)$1.35/1M28.4
08Anthropic: Claude Opus 4.8 (batch)$10.00/1M5.7
09Anthropic: Claude Opus 4.7 (batch)$10.00/1M5.5

Daily tracking began Aug 9- this becomes a live chart once there's a third day to plot.

Index date August 8, 2026Data through 00:00 UTC August 9, 2026Calculated 08:38 UTC

How this is calculated

Quality per dollar= Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). No weighting, no composite - a plain ratio of two real, independently sourced numbers, computed fresh from that day's pricing and benchmark snapshot. The Intelligence Index already measures per-response quality (accuracy across a benchmark suite of reasoning, coding, and knowledge tasks) - it is not response volume, so this ratio isn't "cheap model, more tries." A small quality gap next to a huge price gap is what drives a cheap model to the top.

What this ratio can't tell you:

  • Consistency - whether a model reliably hits the same quality run to run, not just on average.
  • Cost of a wrong answer in your specific use case - a bad response in production code is not the same as a bad response in casual chat.
  • Whether a task needs multiple attempts to reach a usable result.

This is a real, industry-standard way to pre-filter candidates by cost-efficiency - not a final verdict. For high-stakes tasks, weigh the absolute Intelligence Index (shown alongside every ratio below) more heavily than the ratio itself, and check a model's measured endpoint uptime for its consistency track record.

Top 15 models by quality per dollar

Intelligence Index points per dollar of blended price - the 15 best-value models among all 85 with a published quality score. This is a ratio, not a capability ranking: an ultra-cheap model with modest absolute quality can outscore a much stronger, pricier one, purely because price sits near zero. Each bar shows its own absolute Intelligence Index and price alongside the ratio, so you can see which is driving the number.

01Ling-3.0-flash1200.0pts/$Quality 37.8 · $0.03/1M
02inclusionAI: Ling-2.6-flash946.7pts/$Quality 14.2 · $0.02/1M
03Z.ai: GLM 5.2 (batch)489.3pts/$Quality 52.6 · $0.11/1M
04DeepSeek: DeepSeek V4 Flash 0423460.4pts/$Quality 51.8 · $0.11/1M
05Tencent: Hy3 preview423.1pts/$Quality 42.2 · $0.10/1M
06OpenAI: gpt-oss-120b343.1pts/$Quality 24.1 · $0.07/1M
07OpenAI: gpt-oss-20b (free)276.4pts/$Quality 15.2 · $0.05/1M
08OpenAI: GPT-5.6 Luna (batch)232.4pts/$Quality 52.3 · $0.22/1M
09Xiaomi: MiMo-V2.5217.1pts/$Quality 38.0 · $0.18/1M
10Qwen: Qwen3.5-9B193.8pts/$Quality 21.8 · $0.11/1M
11Google: Gemma 4 26B A4B (free)189.8pts/$Quality 26.1 · $0.14/1M
12Google: Gemma 4 31B (free)185.6pts/$Quality 29.7 · $0.16/1M
13NVIDIA: Nemotron 3 Nano 30B A3B (free)165.7pts/$Quality 14.5 · $0.09/1M
14inclusionAI: Ring-2.6-1T146.4pts/$Quality 31.1 · $0.21/1M
15DeepSeek: DeepSeek V3.2108.0pts/$Quality 32.6 · $0.30/1M

Open vs. closed AI, tracked daily

35 days of real usage history

Open-weight models now handle 74% of real token volume on OpenRouter - up 5 points since July 5, 2026. No other site tracks this as a living, dated trend.

100%75%50%25%0%Share of usageJul 5Jul 22Aug 8
  • Open-source
  • Closed

Best AI model for…

Ranked by real usage share for that specific task, not overall popularity.

code

Code Generation

  1. 1. DeepSeek: DeepSeek V4 Flash 042320%
  2. 2. OpenAI: GPT-5.6 Luna (batch)9%
  3. 3. Xiaomi: MiMo-V2.58%

DeepSeek: DeepSeek V4 Flash 0423 leads with 20% of task usage, 10 points ahead of OpenAI: GPT-5.6 Luna (batch).

Full code generation ranking, 9 models

general

Content Writing

  1. 1. DeepSeek: DeepSeek V4 Flash 042322%
  2. 2. Google: Gemini 2.5 Flash Lite (batch)7%
  3. 3. Google: Gemini 2.5 Flash (batch)6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 22% of task usage, 16 points ahead of Google: Gemini 2.5 Flash Lite (batch).

Full content writing ranking, 8 models

code

Debugging

  1. 1. DeepSeek: DeepSeek V4 Flash 042319%
  2. 2. Z.ai: GLM 5.2 (batch)11%
  3. 3. Xiaomi: MiMo-V2.59%

DeepSeek: DeepSeek V4 Flash 0423 leads with 19% of task usage, 8 points ahead of Z.ai: GLM 5.2 (batch).

Full debugging ranking, 9 models

general

Translation

  1. 1. DeepSeek: DeepSeek V4 Flash 042324%
  2. 2. Google: Gemini 2.5 Flash Lite (batch)17%
  3. 3. OpenAI: GPT-4o-mini (batch)8%

DeepSeek: DeepSeek V4 Flash 0423 leads with 24% of task usage, 7 points ahead of Google: Gemini 2.5 Flash Lite (batch).

Full translation ranking, 9 models

general

Summarization

  1. 1. DeepSeek: DeepSeek V4 Flash 042314%
  2. 2. OpenAI: gpt-oss-120b11%
  3. 3. Google: Gemini 2.5 Flash (batch)6%

DeepSeek: DeepSeek V4 Flash 0423 leads with 14% of task usage, 4 points ahead of OpenAI: gpt-oss-120b.

Full summarization ranking, 9 models

general

Customer Support

  1. 1. Google: Gemini 2.5 Flash (batch)20%
  2. 2. DeepSeek: DeepSeek V4 Flash 042313%
  3. 3. Google: Gemini 3 Flash Preview (batch)8%

Google: Gemini 2.5 Flash (batch) leads with 20% of task usage, 7 points ahead of DeepSeek: DeepSeek V4 Flash 0423.

Full customer support ranking, 9 models

general

Roleplay & Fiction

  1. 1. DeepSeek: DeepSeek V4 Flash 042347%
  2. 2. DeepSeek: DeepSeek V4 Pro5%
  3. 3. DeepSeek: DeepSeek V3.24%

DeepSeek: DeepSeek V4 Flash 0423 leads with 47% of task usage, 42 points ahead of DeepSeek: DeepSeek V4 Pro.

Full roleplay & fiction ranking, 9 models

general

Q&A & Knowledge

  1. 1. DeepSeek: DeepSeek V4 Flash 042320%
  2. 2. Google: Gemini 3 Flash Preview (batch)5%
  3. 3. OpenAI: GPT-5.6 Luna (batch)4%

DeepSeek: DeepSeek V4 Flash 0423 leads with 20% of task usage, 16 points ahead of Google: Gemini 3 Flash Preview (batch).

Full q&a & knowledge ranking, 9 models

general

Classification

  1. 1. DeepSeek: DeepSeek V4 Flash 042315%
  2. 2. Google: Gemini 2.5 Flash Lite (batch)10%
  3. 3. OpenAI: GPT-5.6 Luna (batch)7%

DeepSeek: DeepSeek V4 Flash 0423 leads with 15% of task usage, 5 points ahead of Google: Gemini 2.5 Flash Lite (batch).

Full classification ranking, 9 models

Rankings by price and specialty

Comparing a 7B model against a flagship on one scale is meaningless. Each of these is a self-contained "best model" claim within a comparable price tier or use case.

By price

By specialty

Daily snapshots

Every published day gets its own permanent, unchanging permalink - a verifiable point in time for citations, not a page that silently drifts.

How this is calculated

No composite index. No invented weights. Two plain, independently-checkable numbers, each plotted as a time series instead of a single snapshot:

Quality per dollar

Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion). A plain division, recomputed from that day's real pricing and benchmark snapshot.

Usage share

Real token volume through OpenRouter, as a share of that day's total - no sampling, no estimation.

We tried composite indices three times this year (a weighted-sum composite, a weighted-geometric-mean composite, and a usage-vs-quality divergence index) and killed all three after finding real defensibility problems - see the full methodology page for that history. The lesson: publish real numbers over time, not invented ones.

Source: OpenRouter (openrouter.ai/rankings), as of August 9, 2026.

Benchmark scores: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

Methodology: smophy.ai/benchmark/methodology

Token counts originate from each provider's own tokenizer and are not directly comparable across providers.

Journalist or researcher? Pre-computed citable stats and CSV/JSON exports →