The Lab

Model Insights


What the model market buys, what it pays, what it scores, and how much of it you could run yourself. Every source read on a schedule, and kept apart where they disagree.

Swept daily at 6:15 AM Eastern · traffic day Sep 14, 2026 · second gateway through Sep 15, 2026 · series began Aug 24, 2026

18.1Ttokens, one day
805.4Mrequests
486models with traffic
1,283priced offers
350days of daily tape
56%of tokens on open weights

Where the tokens go

One day of traffic across every model the gateway served, re-drawn from each sweep. Concentration is the standing story of this chart: attention is rarely evenly spread, and how top-heavy the day is tells you as much as who is on top. Every model name links to its page on the gateway.

SEP 14, 2026 · 18.1T TOKENS · TOP MODEL 14.2% · TOP FIVE 49.1% · TOP TWENTY 80.7%

Rising and falling

A model’s share of the latest day against its share of the trailing month, both measured by this instrument in the same sweep. A line sloping up means the model took a bigger cut of the day than its month average — attention arriving, not just volume. Models below 100B tokens for the month are excluded; a multiple on a tiny base is noise wearing a percent sign.

TRAILING MONTHSEP 14, 2026inclusionai/ling-3.0-flash-vlinclusionai/ling-3.0-flash-vl0.5% 8.2×deepseek/deepseek-v4.1-flashdeepseek/deepseek-v4.1-flash7.1% 5.4×nex-agi/nex-n2.5-mininex-agi/nex-n2.5-mini0.1% 4.5×nex-agi/nex-n2.5-pronex-agi/nex-n2.5-pro0.3% 3.4×qwen/qwen3.8-flashqwen/qwen3.8-flash0.4% 3.1×openai/gpt-6-astraopenai/gpt-6-astra0.6% 2.8×inclusionai/ling-3.0-flash-santeling-3.0-flash-sante0.3% 2.6×google/gemini-3.7-flashgoogle/gemini-3.7-flash0.5% 0.3×google/gemini-3.6-flashgoogle/gemini-3.6-flash0.1% 0.2×minimax/minimax-m2.7minimax/minimax-m2.70.1% 0.2×meta/muse-spark-1.2-contributormuse-spark-1.2-contributor<0.1% 0.1×stepfun/step-3.7-flashstepfun/step-3.7-flash<0.1% 0.0×

The agentic tell

Tokens divided by requests. A person typing questions spends a few thousand tokens a turn; an agent working through a codebase spends tens of thousands. The ratio separates the two without anyone declaring which they are.

nvidia/nemotron-3-ultra-550b-a55b111,7524.1M req
tencent/hy4-preview110,98415.5M req
tencent/hy378,2059.9M req
z-ai/glm-5.374,4964.0M req
xiaomi/mimo-v2.572,69115.9M req
openai/gpt-5.6-sol65,3294.8M req
deepseek/deepseek-v4.1-flash58,78621.8M req
openai/gpt-5.6-luna41,36962.1M req
z-ai/glm-5.3-flash38,80140.6M req
google/gemini-3.8-flash34,6378.6M req
deepseek/deepseek-v4-flash-073124,04572.3M req
deepseek/deepseek-v4-flash10,09364.2M req

Among the models on the demand board, the busiest by requests in the current reading is not the busiest by tokens: deepseek/deepseek-v4-flash-0731 handled 72.3M requests at 24,045 tokens each, while nvidia/nemotron-3-ultra-550b-a55b ran 111,752 tokens a request — same gateway, two different jobs.

Who is actually spending it

The gateway names 20 applications behind its traffic, ranked by its own numbering — which runs to 26: where the ranks skip a number, the gateway published no application for that slot, so the gap is the source’s, not a missing row. A marks an application whose repository this site’s Repos instrument follows — one instrument watches who builds a thing, the other watches whose tokens it spends, and this column is where they meet. Both are read on the same daily cadence, so a repository that starts drawing tokens today is marked today — and a is one followed daily that has not yet cleared the floors the Repos panel publishes, so there is no page to send you to for it yet.

In the current reading, 16 of the 20 are repositories the Repos instrument follows — hermes-agent, claude-code, kilocode, cline among them, and 13 of the 20 are command-line or IDE agents rather than chat interfaces. 9 of those 16 appear on the published panel today; the rest are followed daily but sit below the floors that panel publishes, and reach it at its next monthly rebuild. When most of the board is agents, that is the same conclusion the tokens-per-request chart reaches from the other direction — and two independent measurements agreeing is worth more than either alone.

The same lab at two doors

Red is the second gateway’s share of tokens; gray is the share computed from this instrument’s own daily sweep of the first. A row reading no join is a name that does not match across the two sources, shown rather than hidden — dropping it would quietly flatter the overlap.

The widest disagreement in the current reading is xiaomi: 0.4% of one gateway’s tokens against 6.6% of the other’s, a 15× difference in how big the same lab looks depending on where you stand.

VERCELOPENROUTERdeepseek66.5%22.8%openai7.6%19.1%anthropic6.9%4.1%zai5.2%11.8%moonshotai3.7%1.4%google3.2%5.6%meta2.6%1.9%stepfun2.3%0.0%spacexaino join0.5%xiaomi0.4%6.6%minimax0.3%1.2%

Neither gateway is the market. Each is a customer base, and the divergence between them measures who routes where — it is not an error to be reconciled away. Neither sees traffic sent straight to a lab’s own API, which is where most frontier revenue lands. The two shares also cut differently: the second gateway’s is its text traffic, while this instrument’s sweep counts everything the first gateway served that day. And a corporate umbrella is folded to the lab that owns it — a model line published under its own name counts toward its parent on both sides.

The tape, day by day

A second gateway publishes a fixed range of daily history, which is the one thing this instrument can go back and get rather than wait for. Share of tokens above, share of spend below, per lab, every day of the range. The lab leading the latest day is drawn in red so the crossings stay readable; the rest are gray on purpose.

SHARE OF TOKENS · 350 DAYS · OCT 1, 2025 SEP 15, 2026

0%19%38%57%76%deepseekopenaianthropiczaimoonshotaigooglemetaOct 1, 2025Sep 15, 2026

SHARE OF SPEND · SAME RANGE

0%25%49%74%98%anthropicopenaimoonshotaideepseekgooglezaispacexaiOct 1, 2025Sep 15, 2026

The daily sweep the rest of these readings rest on is forward-only: its source serves a rolling window and sells no history at any price, so its record begins the day collection began. This range is re-read whole every morning and arrives in a single request. The two are not interchangeable, and that difference is why both are here. Each chart names the seven largest labs by its own metric on the latest day, so the two lists can differ.

Volume against money

Every lab on one line: where its dots sit is its share of tokens against its share of spend, and the bar between them is the finding. A lab whose red dot sits far to the right of its gray one is being paid a premium for every token it serves.

On Sep 15, 2026 the biggest source of tokens was deepseek at 66.5% of volume, while the biggest source of money was anthropic at 39.4% of spend on 6.9% of the tokens. Per token routed, anthropic collects 5.67× its share of the bill and deepseek collects 0.13× — a 43× difference in what a token is worth depending on whose it is.

SHARE OF TOKENSSHARE OF SPENDanthropicopenaimoonshotaideepseekgooglezaispacexaistepfunnvidiametaalibabaminimax0%36%72%

The same model is not one price

A model is served by many providers at once, and they do not agree. Each bar spans the cheapest and dearest dollars-per-million-prompt-tokens across every provider serving that exact model. Free and disabled routes are excluded — a $0 route is a promotion, not a price.

$0.1$0.5$1$5$ / M PROMPT TOKENSdeepseek-v4-flash-073111.0×glm-5.3-flash6.0×gpt-5.6-sol5.5×glm-5.24.7×deepseek-v4-flash4.2×gpt-5.6-luna4.0×mimo-v2.53.4×minimax-m33.3×kimi-k32.9×deepseek-v4.1-flash2.5×deepseek-v4-pro2.4×glm-5.32.4×

Widest spread in the current reading: deepseek/deepseek-v3.2 spans $0.21 to $3.00 across 15 providers — 14.4× for identical weights, gmicloud cheapest and mara dearest.

And it is not only which company — which tier

openai/gpt-oss-120b · every deepinfra offer on the gateway, Sep 15, 2026

One vendor, identical model — and a 5.4× spread on service tier alone in the current reading. This distinction is the reason the sensor records every offer rather than a headline price. The exemplar above is picked by rule: whichever single provider shows the widest tier spread on the day, not an editor’s favorite.

The price of privacy

Suppose your prompts must not be kept. A provider counts as clean here only if it publishes a data policy declaring it neither trains on nor retains what you send — so each model’s real price becomes the cheapest clean offer, not the cheapest offer. For the highest-scoring models: the filled dot is the cheapest route from anyone, the hollow dot the cheapest clean route, and the gap between them is what privacy costs that day.

$0.1$0.5$1$5$10$ / M PROMPT TOKENSclaude-fable-5.1paritygpt-6-astra2.0× for privacyclaude-opus-5parityclaude-fable-5paritygpt-5.6-sol4.4× for privacyglm-5.3paritygrok-4.61.1× for privacykimi-k3paritygpt-5.6-terra2.0× for privacyclaude-opus-4.8parityglm-5.3-flashparitygemini-3.8-flashparity

In the current reading, privacy costs nothing for 39 of the 60 scored models with a clean route — the clean provider already is the cheapest.

What a point of intelligence costs

Price alone says nothing without a denominator. Each model’s cheapest real price is plotted against an independent capability score (Artificial Analysis’s intelligence index) — the down-and-right edge is the efficiency frontier, and over a series this chart becomes the answer to the only question that matters about inference cost: is that frontier moving, and how fast.

$0.1$0.5$1$5$10CHEAPEST $ / M PROMPT TOKENS →INTELLIGENCE INDEX →deepseek-v4-flash-0731claude-opus-5gpt-5.6-solglm-5.3-flashclaude-sonnet-5glm-5.2nemotron-3-ultra-550b-a55bgpt-oss-120bclaude-sonnet-4.5glm-5.3gpt-6-astragpt-5.6-lunadeepseek-v4.1-flashmimo-v2.5deepseek-v4-flashgemini-3.8-flash

The cheapest point of capability in the current reading is deepseek/deepseek-v4-flash-0731 at $0.0012 per index point; the highest-scoring model, anthropic/claude-fable-5.1, costs $0.187 per point — a 156× difference in what a unit of measured capability costs.

The same chart, by the job

“Smartest per dollar” is rarely the actual question — a developer wants the smartest at their job per dollar. The same frontier, re-plotted against the independent coding and agentic scores. The solid mark is the best value on each.

Coding index

$0.1$1$10deepseek-v4-flash-0731claude-sonnet-5gpt-5.6-solnemotron-3-ultra-550b-a55bmimo-v2.5claude-sonnet-4.5gpt-oss-120bgemini-2.5-procommand-adeepseek-v4-flashgemini-3.8-flashglm-5.3CHEAPEST $ / M →

Agentic index

$0.1$1$10deepseek-v4-flash-0731claude-sonnet-5gpt-5.6-solnemotron-3-ultra-550b-a55bclaude-sonnet-4.6claude-haiku-4.5deepseek-v4-flashmimo-v2.5claude-sonnet-4.5glm-5.3-flashgemini-3.8-flashglm-5.3CHEAPEST $ / M →

In the current reading, one model is the value pick for both jobs: deepseek/deepseek-v4-flash-0731.

The speed frontier

For anything interactive, tokens per second matters as much as the score. Best measured throughput against intelligence; the labeled edge is the set of models where nothing faster is smarter. Throughput is measured on each model’s default route only, so a model can run faster than shown from a different provider — never slower than a measurement that exists.

MEDIAN THROUGHPUT, TOKENS / SECOND →INTELLIGENCE INDEX →50100150200claude-opus-5gemini-3.8-flashgpt-5.6-lunagemini-3.7-flashmuse-spark-1.1mimo-v2.5gemini-3.5-flash-litemercury-2claude-fable-5.1deepseek-v4-flash-0731glm-5.3-flashdeepseek-v4.1-flashdeepseek-v4-flashgpt-5.6-solglm-5.3minimax-m3

The fastest model in the smarter half of the field, in the current reading: meta/muse-spark-1.1 at 197 tokens per second.

What a window costs

The cheapest priced route at each context-window tier — for when the question is not how smart the model is, but how much it can hold.

The free tier

Free is technically the best price, so price stops being an axis. What replaces it is what the zero costs: every scored model the gateway serves at $0, plotted by intelligence against the worst measured string its serving providers carry — from no declared strings to training on the prompts you send. The solid mark is the highest score in the cleanest occupied column.

“Free” has a price, and the sensor measures it. In the current reading, 5 of the 7 scored free models are served by a provider that trains on your prompts, and 6 of the 7 by one that retains them. Free routes are promotions, not commitments — they can be withdrawn, rate-limited, or deprioritized without notice. And a pre-release running under a placeholder label carries no independent score, so it cannot appear here no matter how capable it looks.

NO DECLARED STRINGSRETAINS PROMPTSNO POLICYTRAINS ON PROMPTSFEWER STRINGS ←INTELLIGENCE INDEX →inkling-small:freeinkling:freeling-3.0-flash-vl:freenemotron-3-ultra-550b-a55b:freegemma-4-31b-it:freenemotron-3-super-120b-a12b:free

Can you actually download the weights?

Every model carrying traffic on the daily sweep is probed against the public model hub. The test is not whether a repository exists with a matching name. It is whether that repository contains downloadable weight files, published by the model’s own author. A model the probe has not yet reached is counted in its own row, never assumed.

TRAFFIC DAY SEP 14, 2026 · WEIGHTS PROBED SEP 15, 2026

Evidence tierModelsShare of tokens
1 · author-published weights15056.4%
2 · weights found, but not from the author146.7%
3 · the author’s own repo, no weights in it70.1%
4 · someone else’s repo, no weights2318.8%
5 · no weights repo found28718.1%
6 · not yet probed this sweep40.0%

56.4% of the day’s tokens ran on models whose authors publish downloadable weights, across 150 models. Only that top tier is a claim. The rest are degrees of not-knowing, kept visible rather than rounded into the headline — a model with no weights repository found is recorded as not found, never as closed, and a model the probe list has not yet reached is counted in the last row rather than folded into any verdict.

The first version of this test matched on name alone and reported more than double whatever the top tier reads. Matching by name declared closed frontier models open, because someone had registered a repository wearing each name with nothing inside it.

Repositories wearing a name

Some probed models match a public repository belonging to someone other than the model’s author, whose name normalizes to theirs, and which contains no weights at all. They are kept rather than discarded: a repository carrying a closed frontier model’s name is a finding in its own right.

In the current reading that is 23 models, carrying 18.8% of the day’s tokens by the traffic of the models being impersonated; the table shows the 14 busiest by that traffic. A model’s own author publishing a card-only repository is a different thing entirely, and is counted in its own tier above rather than here.

ModelRepository claiming the nameFilesDownloads
openai/gpt-5.6-lunacrosbylegal/gpt-5.6-luna30
openai/gpt-5.6-solincoggamer/gpt-5.6-sol10
anthropic/claude-sonnet-5incoggamer/claude-sonnet-510
openai/gpt-5.6-terracrosbylegal/gpt-5.6-terra30
google/gemini-2.5-flashMrRobGappy/Gemini_2.5_flash10
openai/gpt-4o-minimartianband1t/gpt4omini20
qwen/qwen3.7-plusRscriptSQwen/Qwen3.7-plus20
google/gemini-3.6-flashcrosbylegal/gemini-3.6-flash30
openai/gpt-4.1-miniLogambika/gpt-4.1-mini20
openai/gpt-5-nanoCometAPI/gpt-5-nano20
x-ai/grok-4.3incoggamer/grok-4.310
google/gemini-2.5-proafu4642tD/gemini-2.5-pro10
anthropic/claude-sonnet-4CometAPI/Claude_Sonnet420
google/gemini-2.5-flash-imageCometAPI/gemini-2.5-flash-image20

And the harder case — 14 models whose weights were found, but not published by the model’s author; the table shows the 8 most-downloaded of them. Some are genuine, an organization whose name simply differs. Some are unrelated models that happen to share a word. This is why author-verification, and not the mere presence of weights, is what the headline is allowed to rest on.

ModelRepository foundWeight filesDownloads
bytedance/ui-tars-1.5-7bByteDance-Seed/UI-TARS-1.5-7B7612,032
cognitivecomputations/dolphin-mistral-24b-venice-editiondphn/Dolphin-Mistral-24B-Venice-Edition10470,367
meta/muse-glimmer-30bmeta-models/Muse-Glimmer-30B2441,881
xiaomi/mimo-v2.5XiaomiMiMo/MiMo-V2.518307,960
xiaomi/mimo-v2.5-proXiaomiMiMo/MiMo-V2.5-Pro3424,568
microsoft/wizardlm-2-8x22balpindale/WizardLM-2-8x22B599,983
nvidia/nemotron-3-nano-30b-a3bunsloth/Nemotron-3-Nano-30B-A3B131,380
meituan/longcat-2.0meituan-longcat/LongCat-2.01941,367

Where open-weight models come from

Open-weight releases per year among notable models, US-affiliated against China-affiliated. Affiliation is recorded per author, so a model with authors in both places counts for both — these are two counts, not a split of one.

Across the complete years, US-affiliated open-weight releases peaked at 37 in 2023 and stood at 11 in 2025, the last complete year. Chinese-affiliated releases went from 13 to 28 across the same span. China first published more open-weight notable models than the US in 2025 (28 to 11).

US-AFFILIATEDCHINA-AFFILIATED271201921420203110202132720223713202320162024112820254172026

From a curated registry of 1,065 notable models, 814 of which carry an explicit open-weights verdict. It is deep on provenance and thin on the live commercial frontier — a model absent from it is outside its notability criteria, not unreal. The final year on the chart is partial by definition: it is still being published, which is why every sentence above stops at the last complete year.

What a parameter costs

Only author-published open-weight models appear here: they are the only ones whose true size is knowable. Circle area is that model’s measured token volume and both axes are logarithmic. Labels go to the largest by traffic, and a label that would collide is dropped rather than moved.

The cheapest route on the board in the current reading is meta-llama/llama-3.1-8b-instruct at $0.04 per million output tokens for 8.0B parameters. The largest is moonshotai/kimi-k3 at 2780B and $10.95. Size and price are related but not tightly, and the spread across the middle of this chart is the part a headline price cannot show.

$0.1$1$101B10B100B1000Bdeepseek/deepseek-v4-flash-0731tencent/hy4-previewz-ai/glm-5.3-flashdeepseek/deepseek-v4.1-flashtencent/hy3deepseek/deepseek-v4-flashz-ai/glm-5.2minimax/minimax-m3$ / MILLION OUTPUT TOKENSPARAMETERS

The parameter counts come from each repository’s own weight index. The gateway catalog carries a field that looks like a parameter count and is not one — its values sit between 6 and 19 across nearly every model, and anything sized from it would be wrong by nine orders of magnitude.

Where the compute sits

Provider headquarters

US55
Undisclosed32
SG6
CN6
IL2
NL1
ID1
SE1
ES1
FR1

Declared data policy

85 of 106 providers publish one.

4 providers train on prompts they are sent.

33 retain prompts without training on them.

21 publish no policy at all — recorded as unknown, never as a no.

That last line is a measurement decision, not a footnote. A missing policy and a policy that says “we do not train” are different facts, and collapsing them would invent the most consequential number on this page.

Reliability varies more than reputation suggests: of 1,304 offers, 697 held above 99.5% uptime over 24 hours while 33 sat below 80%. Uptime is a live reading — it cannot be reconstructed afterwards, which is why it is taken daily.

What these readings cannot tell you

Long trends. The series began Aug 24, 2026. Rank changes and the shape of the demand curve sharpen with every sweep; the trailing-month figures above come from the gateway’s own 30-day aggregates, recorded the same way daily.

The full application picture. The gateway publishes only 20 named applications — a leaderboard, not a census, and one with unnamed slots in its own ranking. The tail is invisible.

Anything about tool use. The source exposes tool-call and cached-token fields that are present on every row and — measured this sweep — zero on all of them. They are placeholders, and nothing on this page is claimed from them; the collector is set to announce the day they start carrying data.

The market beyond this gateway. Traffic routed directly to a vendor never appears here. What the gateway shows is the open, price-comparing slice of the market — the slice where competition is visible.

Source: OpenRouter’s public catalog and rankings, read once a day by a collector this instrument runs on its own hardware. The gateway publishes no history — every number here exists because it was recorded on the day it was true, and a missed day cannot be bought back at any price. Figures describe traffic routed through this one gateway; traffic sent directly to a vendor never appears in it.

Coverage, measured

Coverage differs by dataset and by metric, so it is counted from the stored rows rather than asserted. Two limits are worth stating plainly next to it: the second gateway publishes shares of its own traffic and never volumes, so a share here cannot be multiplied by a token count from anywhere else and two gateways’ shares cannot be added; and the join between sources is a lossy name normalization, so where it fails the row is shown unjoined rather than dropped.

GroupMetricDaysFirstLast
labrequests3502025-10-012026-09-15
labspend3502025-10-012026-09-15
labtokens3502025-10-012026-09-15
modelrequests3502025-10-012026-09-15
modelspend3362025-10-152026-09-15
modeltokens1452026-04-242026-09-15

Sources and licenses

  • epoch-notable-ai-models1,065 rows, read through Sep 15, 2026, CC-BY-4.0
  • huggingface-hub-api481 rows, read through Sep 15, 2026, no separate data license asserted
  • vercel-leaderboard-export108,359 rows, read through Sep 15, 2026, CC-BY-4.0

AI Gateway Leaderboard Data © Vercel, used under CC BY 4.0. Data on Notable AI Models © Epoch AI, used under CC BY 4.0. Model metadata from the Hugging Face Hub. Model volume, pricing, and endpoint data collected by this instrument.