LLM Rankings

The models that actually carry the load.

Benchmarks measure potential. Token volume measures trust. Track the top 20 models by real production usage — with intelligence, coding, and GPQA scores alongside — updated daily.

Updated Sep 8, 2026
Tokens in window
425T
30 periods of traffic
Top model share
11.8%
DeepSeek: DeepSeek V4 Flash 0731 (batch)
Top 50 share
94.2%
of all OpenRouter traffic
Models ranked
20
2026-08-09 → 2026-09-07
Token volume over time
09.53T19.1TAug 9 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.33T tokensAug 9 — OpenAI: GPT-5.6 Luna (batch): 563B tokensAug 9 — Tencent: Hy3: 1.55T tokensAug 9 — Other models: 5.53T tokensAug 9Aug 10 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.58T tokensAug 10 — OpenAI: GPT-5.6 Luna (batch): 789B tokensAug 10 — Tencent: Hy3: 1.77T tokensAug 10 — Other models: 6.48T tokensAug 11 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.76T tokensAug 11 — OpenAI: GPT-5.6 Luna (batch): 815B tokensAug 11 — Tencent: Hy3: 1.51T tokensAug 11 — Other models: 6.84T tokensAug 12 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.72T tokensAug 12 — OpenAI: GPT-5.6 Luna (batch): 871B tokensAug 12 — Tencent: Hy3: 1.58T tokensAug 12 — Other models: 7.06T tokensAug 13 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.63T tokensAug 13 — OpenAI: GPT-5.6 Luna (batch): 790B tokensAug 13 — Tencent: Hy3: 1.50T tokensAug 13 — Other models: 7.21T tokensAug 14 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.62T tokensAug 14 — OpenAI: GPT-5.6 Luna (batch): 746B tokensAug 14 — Tencent: Hy3: 1.42T tokensAug 14 — Other models: 7.70T tokensAug 14Aug 15 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.41T tokensAug 15 — OpenAI: GPT-5.6 Luna (batch): 652B tokensAug 15 — Tencent: Hy3: 1.12T tokensAug 15 — Other models: 6.39T tokensAug 16 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.46T tokensAug 16 — OpenAI: GPT-5.6 Luna (batch): 648B tokensAug 16 — Tencent: Hy3: 1.06T tokensAug 16 — Other models: 7.21T tokensAug 17 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.69T tokensAug 17 — OpenAI: GPT-5.6 Luna (batch): 1.05T tokensAug 17 — Tencent: Hy3: 1.52T tokensAug 17 — Other models: 9.06T tokensAug 18 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.74T tokensAug 18 — OpenAI: GPT-5.6 Luna (batch): 1.04T tokensAug 18 — Tencent: Hy3: 1.63T tokensAug 18 — Other models: 8.29T tokensAug 19 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.68T tokensAug 19 — OpenAI: GPT-5.6 Luna (batch): 807B tokensAug 19 — Tencent: Hy3: 1.22T tokensAug 19 — Other models: 8.41T tokensAug 19Aug 20 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.83T tokensAug 20 — OpenAI: GPT-5.6 Luna (batch): 656B tokensAug 20 — Tencent: Hy3: 1.25T tokensAug 20 — Other models: 8.39T tokensAug 21 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.85T tokensAug 21 — OpenAI: GPT-5.6 Luna (batch): 555B tokensAug 21 — Tencent: Hy3: 1.12T tokensAug 21 — Ox Alpha: 1.99T tokensAug 21 — Other models: 8.37T tokensAug 22 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.39T tokensAug 22 — OpenAI: GPT-5.6 Luna (batch): 417B tokensAug 22 — Tencent: Hy3: 754B tokensAug 22 — Ox Alpha: 4.54T tokensAug 22 — Other models: 6.97T tokensAug 23 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.39T tokensAug 23 — OpenAI: GPT-5.6 Luna (batch): 386B tokensAug 23 — Tencent: Hy3: 707B tokensAug 23 — Ox Alpha: 5.02T tokensAug 23 — Other models: 7.64T tokensAug 24 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.71T tokensAug 24 — OpenAI: GPT-5.6 Luna (batch): 613B tokensAug 24 — Tencent: Hy3: 1.07T tokensAug 24 — Ox Alpha: 5.93T tokensAug 24 — Other models: 9.03T tokensAug 24Aug 25 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.77T tokensAug 25 — OpenAI: GPT-5.6 Luna (batch): 724B tokensAug 25 — Tencent: Hy3: 1.03T tokensAug 25 — Ox Alpha: 5.75T tokensAug 25 — Other models: 9.50T tokensAug 26 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 2.41T tokensAug 26 — OpenAI: GPT-5.6 Luna (batch): 753B tokensAug 26 — Tencent: Hy3: 1.04T tokensAug 26 — Ox Alpha: 4.00T tokensAug 26 — Other models: 10.2T tokensAug 27 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 2.00T tokensAug 27 — OpenAI: GPT-5.6 Luna (batch): 1.45T tokensAug 27 — Tencent: Hy3: 1.11T tokensAug 27 — Other models: 11.2T tokensAug 28 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.61T tokensAug 28 — OpenAI: GPT-5.6 Luna (batch): 1.70T tokensAug 28 — Tencent: Hy3: 963B tokensAug 28 — Other models: 10.8T tokensAug 29 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.37T tokensAug 29 — OpenAI: GPT-5.6 Luna (batch): 1.12T tokensAug 29 — Tencent: Hy3: 699B tokensAug 29 — Other models: 10.0T tokensAug 29Aug 30 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.44T tokensAug 30 — OpenAI: GPT-5.6 Luna (batch): 1.43T tokensAug 30 — Tencent: Hy3: 748B tokensAug 30 — Other models: 9.70T tokensAug 31 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.59T tokensAug 31 — OpenAI: GPT-5.6 Luna (batch): 1.30T tokensAug 31 — Tencent: Hy3: 825B tokensAug 31 — Other models: 10.6T tokensSep 1 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.63T tokensSep 1 — OpenAI: GPT-5.6 Luna (batch): 1.76T tokensSep 1 — Tencent: Hy3: 515B tokensSep 1 — Other models: 12.3T tokensSep 2 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.73T tokensSep 2 — OpenAI: GPT-5.6 Luna (batch): 2.84T tokensSep 2 — Tencent: Hy3: 730B tokensSep 2 — Other models: 13.2T tokensSep 3 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.91T tokensSep 3 — OpenAI: GPT-5.6 Luna (batch): 1.44T tokensSep 3 — Tencent: Hy3: 768B tokensSep 3 — Other models: 14.4T tokensSep 3Sep 4 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 2.24T tokensSep 4 — OpenAI: GPT-5.6 Luna (batch): 1.39T tokensSep 4 — Tencent: Hy3: 510B tokensSep 4 — Other models: 14.1T tokensSep 5 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.73T tokensSep 5 — OpenAI: GPT-5.6 Luna (batch): 2.01T tokensSep 5 — Tencent: Hy3: 343B tokensSep 5 — Other models: 11.3T tokensSep 6 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.52T tokensSep 6 — OpenAI: GPT-5.6 Luna (batch): 2.19T tokensSep 6 — Tencent: Hy3: 301B tokensSep 6 — Other models: 10.3T tokensSep 7 — DeepSeek: DeepSeek V4 Flash 0731 (batch): 1.58T tokensSep 7 — OpenAI: GPT-5.6 Luna (batch): 2.56T tokensSep 7 — Tencent: Hy3: 495B tokensSep 7 — Other models: 14.4T tokens
DeepSeek: DeepSeek V4 Flash 0731 (batch)OpenAI: GPT-5.6 Luna (batch)Tencent: Hy3Ox AlphaOther models
Token volume by period and model
PeriodDeepSeek: DeepSeek V4 Flash 0731 (batch)OpenAI: GPT-5.6 Luna (batch)Tencent: Hy3Ox AlphaOther models
2026-08-091.33T563B1.55T05.53T
2026-08-101.58T789B1.77T06.48T
2026-08-111.76T815B1.51T06.84T
2026-08-121.72T871B1.58T07.06T
2026-08-131.63T790B1.50T07.21T
2026-08-141.62T746B1.42T07.70T
2026-08-151.41T652B1.12T06.39T
2026-08-161.46T648B1.06T07.21T
2026-08-171.69T1.05T1.52T09.06T
2026-08-181.74T1.04T1.63T08.29T
2026-08-191.68T807B1.22T08.41T
2026-08-201.83T656B1.25T08.39T
2026-08-211.85T555B1.12T1.99T8.37T
2026-08-221.39T417B754B4.54T6.97T
2026-08-231.39T386B707B5.02T7.64T
2026-08-241.71T613B1.07T5.93T9.03T
2026-08-251.77T724B1.03T5.75T9.50T
2026-08-262.41T753B1.04T4.00T10.2T
2026-08-272.00T1.45T1.11T011.2T
2026-08-281.61T1.70T963B010.8T
2026-08-291.37T1.12T699B010.0T
2026-08-301.44T1.43T748B09.70T
2026-08-311.59T1.30T825B010.6T
2026-09-011.63T1.76T515B012.3T
2026-09-021.73T2.84T730B013.2T
2026-09-031.91T1.44T768B014.4T
2026-09-042.24T1.39T510B014.1T
2026-09-051.73T2.01T343B011.3T
2026-09-061.52T2.19T301B010.3T
2026-09-071.58T2.56T495B014.4T
Model leaderboard● Live
Large language models ranked by tokens processed on OpenRouter over the last 30 days, with benchmark scores where available.
#ModelTokens
1
DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek
11.8% share · $0.14/M in · 1049K ctx
50.3T
2
OpenAI: GPT-5.6 Luna (batch)openai
8.0% share · $0.10/M in · 1050K ctx
34.1T
3
Tencent: Hy3tencent
7.3% share · $0.13/M in · 262K ctx
30.9T
4
Ox Alphastealth
6.4% share
27.2T
5
Xiaomi: MiMo-V2.5xiaomi
6.2% share · $0.14/M in · 1050K ctx
26.6T
6
DeepSeek: DeepSeek V4 Flash 0423deepseek
5.2% share · $0.09/M in · 1049K ctx
22.0T
7
Tencent: Hy4 previewtencent
5.0% share · $0.83/M in · 1049K ctx
21.1T
8
Z.ai: GLM 5.3 Flash (batch)z ai
4.8% share · $0.15/M in · 1049K ctx
20.4T
9
Nemotron 3 Ultra 550b A55b 20260604:Freenvidia
3.9% share
16.8T
10
Z.ai: GLM 5.2z ai
3.3% share · $0.97/M in · 1049K ctx
13.9T
11
Google: Gemini 3.7 Flash (batch)google
2.0% share · $0.38/M in · 1049K ctx
8.35T
12
Minimax M3 20260531:Freeminimax
1.9% share
8.17T
13
DeepSeek: DeepSeek V4 Pro 0423deepseek
1.9% share · $0.96/M in · 1049K ctx
8.08T
14
Claude Opus 5 (batch)anthropic
1.8% share · $2.50/M in · 1000K ctx
7.85T
15
MoonshotAI: Kimi K3 (batch)moonshotai
1.6% share · $3.00/M in · 1049K ctx
6.70T
16
MiniMax: MiniMax M3 (batch)minimax
1.6% share · $0.30/M in · 524K ctx
6.63T
17
Laguna S 2.1 20260720:Freepoolside
1.5% share
6.34T
18
OpenAI: GPT-5.6 Sol (batch)openai
1.4% share · $1.00/M in · 1050K ctx
6.00T
19
Z.ai: GLM 5.3z ai
1.3% share · $1.40/M in · 1311K ctx
5.38T
20
Anthropic: Claude Sonnet 5 (batch)anthropic
1.2% share · $1.00/M in · 1000K ctx
4.98T

All models outside the top 50 accounted for 24.5T tokens in this window.

Intelligence Index

Artificial Analysis composite score, all evaluated models

Artificial Analysis composite score, all evaluated models
1Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
53.4
2Qwen3.8 Max
53.4
3GPT-6 Astra (max)
52.8
4Claude Opus 5 (Adaptive Reasoning, Max Effort)
50.7
5Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
49.7
6GPT-5.6 Sol (max)
47.1
7GLM-5.3 (max)
44.9
8Grok 4.6 (high)
44.4
9Kimi K3 (max)
43.8
10GPT-5.6 Terra (max)
42.3
GPQA Diamond

OpenRouter's own eval — graduate-level science questions

OpenRouter's own eval — graduate-level science questions
1Google: Gemini 3.1 Pro Preview
94.4%
2OpenAI: GPT-6 Astra
94.4%
3Google: Gemini 3.7 Flash
94.3%
4OpenAI: GPT-5.5
93.8%
5OpenAI: GPT-5.6 Sol Pro
93.8%
6SpaceXAI: Grok 4.6
93.3%
7Google: Gemini 3.6 Flash
92.8%
8Google: Gemini 3.5 Flash
92.8%
9OpenAI: GPT-5.6 Sol
91.9%
10MoonshotAI: Kimi K3
91.5%

Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings).

The Mr. Bright Role Index

Which model should do which job?

Leaderboards rank models against each other. Businesses hire them for specific jobs. We score every benchmarked model against the marketing roles we actually staff — weighting the signals that predict performance in that seat, not overall cleverness. Weights are published below each role so you can check the maths.

Software Engineer

Claude Opus 5 (batch)95.8

anthropic · $2.50/M

  • 2. OpenAI: GPT-5.6 Sol (batch)91.4
  • 3. MoonshotAI: Kimi K3 (batch)86.9

Code correctness dominates; tool use matters because the job is multi-step.

Coding 45%Agentic 25%Intelligence 20%GPQA 10%
Hire this role →

Web Designer & CRO

Claude Opus 5 (batch)100

anthropic · $2.50/M

  • 2. OpenAI: GPT-5.6 Sol (batch)90.3
  • 3. Z.ai: GLM 5.386.4

Ships front-end code and reasons about experiments, so coding and reasoning share the load.

Coding 35%Intelligence 35%Agentic 30%
Hire this role →

Paid Social Buyer

Z.ai: GLM 5.3 Flash (batch)71.4

z ai · $0.150/M

  • 2. OpenAI: GPT-5.6 Luna (batch)69.8
  • 3. Claude Opus 5 (batch)65.2

Runs long tool-driven loops over ad platforms every day, so agentic ability and cost per run lead.

Agentic 35%Intelligence 30%Cost efficiency 35%
Hire this role →

Search & PPC Manager

Z.ai: GLM 5.3 Flash (batch)70.5

z ai · $0.150/M

  • 2. OpenAI: GPT-5.6 Luna (batch)69.2
  • 3. Claude Opus 5 (batch)65.2

Same daily tool loop as paid social, with more numeric reasoning over search terms.

Agentic 30%Intelligence 35%Cost efficiency 35%
Hire this role →

SEO & AI Search

OpenAI: GPT-5.6 Sol (batch)83.7

openai · $1.00/M

  • 2. Google: Gemini 3.7 Flash (batch)78.8
  • 3. MoonshotAI: Kimi K3 (batch)78.8

Reads large corpora and must be factually tight, so context length and hard-reasoning accuracy carry weight.

Intelligence 35%GPQA 35%Context length 30%
Hire this role →

Lifecycle Marketer

OpenAI: GPT-5.6 Luna (batch)70.3

openai · $0.100/M

  • 2. Z.ai: GLM 5.3 Flash (batch)69.1
  • 3. Claude Opus 5 (batch)60.2

High message volume makes cost per token decisive alongside writing quality.

Intelligence 35%Cost efficiency 40%Agentic 25%
Hire this role →

Creative Director

OpenAI: GPT-5.6 Luna (batch)70.6

openai · $0.100/M

  • 2. Z.ai: GLM 5.3 Flash (batch)67.6
  • 3. Z.ai: GLM 5.362.9

Generates constantly and holds long brand context, so throughput economics and context matter as much as raw capability.

Intelligence 40%Cost efficiency 30%Context length 30%
Hire this role →

Community Manager

OpenAI: GPT-5.6 Luna (batch)68.3

openai · $0.100/M

  • 2. Z.ai: GLM 5.3 Flash (batch)67
  • 3. OpenAI: GPT-5.6 Sol (batch)62.5

Answers the public all day: cheap, accurate and steady beats brilliant and expensive.

Intelligence 30%GPQA 20%Cost efficiency 30%Agentic 20%
Hire this role →

Manager & Orchestrator

Claude Opus 5 (batch)92.6

anthropic · $2.50/M

  • 2. Z.ai: GLM 5.389.3
  • 3. OpenAI: GPT-5.6 Sol (batch)83.7

Reviews other agents' work and coordinates them, which is the hardest reasoning and longest context job on the team.

Intelligence 40%Agentic 35%Context length 25%
Hire this role →

How the index is calculated

Every signal is min-max normalised to 0–100 across the benchmarked models in the current window, then combined using the weights shown on each card. Cost efficiency is the inverse of prompt price; context is log-scaled so a 1M-token window doesn't swamp every other signal. A model is only ranked for a role when at least half that role's weight is backed by published data, and scores are re-based on the weight actually available — so a model isn't penalised for a benchmark nobody has run on it. Models without an intelligence score are excluded entirely.

This index is our own analysis built on public data. It is a starting point for choosing a brain, not a guarantee of performance on your workload — the right model for your account is the one that wins your own evaluation.

Usage data: OpenRouter public rankings dataset. Token totals are prompt + completion across all public traffic. Mr. Bright AI is not affiliated with OpenRouter.

How these rankings work

Usage, not hype

Rankings are ordered by real tokens processed across OpenRouter's public traffic — prompt plus completion — not by marketing claims or vibes.

Benchmarks alongside

Each model carries its Artificial Analysis Intelligence and Coding indexes plus OpenRouter's own GPQA Diamond eval, so you can weigh capability against adoption.

Momentum tracked

The growth column compares the second half of the window against the first, so you can spot models being adopted — or abandoned — before the rank flips.

Refreshed daily

The usage dataset rebuilds every day; benchmark indexes update as sources publish new runs. No stale screenshots.

We run on the models at the top of this board.

Every Mr. Bright AI team is model-agnostic — we route each task to whichever model is winning at it right now, and we re-check this data continuously.

See our full stack