SmophyAI

Published daily · Next update 02:00 UTC

Methodology

No invented index. Real numbers, tracked over time.

Every figure published on this site is either a real, independently-sourced quantity (a price, a benchmark score, a usage share, an uptime percentage) or a plain, transparent ratio of two such quantities - quality ÷ price, tokens ÷ context window. There is no weighted composite score anywhere on the site.

Why there's no composite index (the honest version)

This site tried three composite indices before settling on the current approach, and each was killed after a real-data audit, not kept out of caution:

  1. A weighted linear sum of Quality/Economy/Adoption/Readiness - killed because a linear sum lets a model buy a high score with one strong component regardless of how bad another is.
  2. A weighted geometric mean of the same four components - fixed the compensation problem, but its dominant component (Quality) turned out to be little more than a normalized restatement of Artificial Analysis's own published intelligence index, so the composite wasn't structurally unique versus a benchmark that already exists.
  3. A usage-vs-quality divergence index ("usage rank minus quality rank") - genuinely unique (no competitor fuses real usage data with real benchmark data this way), but a real-data test showed it couldn't distinguish a genuinely overlooked model from an expensive one that's rationally avoided on price. Fixing that confound would have meant modeling "expected adoption" from price and availability - reintroducing the exact unauditable, invented-weight problem the first two attempts already failed on.

The pattern across all three: every fix pushed further from "real numbers, sorted" toward "our opinion, disguised as math." That's the lesson this site now runs on.

What we publish instead

Quality per dollar, over time
Artificial Analysis Intelligence Index ÷ blended price per 1M tokens (75% prompt / 25% completion), recomputed daily and plotted as a real time series - not a static snapshot like every competing chart in this space.
Open vs. closed usage share
Real OpenRouter token volume, split by whether a model's weights are open or closed (classified by publisher - see server/weights-class.ts), tracked daily.
Provider reliability
Real per-provider uptime (trailing 30 minutes) and latency, per model, tracked daily - nobody else publishes this as a historical series.
Real throughput
Tokens/second from real provider endpoint data, tracked daily rather than benchmarked once.
Task usage vs. benchmark score (coding and agentic)
Real per-task usage share from OpenRouter's classified-traffic data, shown alongside (never blended with) the relevant Artificial Analysis domain index.

Data freshness & new models

A daily job at 02:00 UTC re-fetches OpenRouter's models list, usage rankings, benchmark scores, provider endpoints, task classifications, and app rankings - and discovers newly released models automatically before syncing the rest, so a brand-new model is picked up without any manual step. Every response is archived before parsing.