The models that actually carry the load.
Benchmarks measure potential. Token volume measures trust. Track the top 20 models by real production usage — with intelligence, coding, and GPQA scores alongside — updated daily.
| Period | DeepSeek: DeepSeek V4 Flash 0731 (batch) | OpenAI: GPT-5.6 Luna (batch) | Tencent: Hy3 | Ox Alpha | Other models |
|---|---|---|---|---|---|
| 2026-08-09 | 1.33T | 563B | 1.55T | 0 | 5.53T |
| 2026-08-10 | 1.58T | 789B | 1.77T | 0 | 6.48T |
| 2026-08-11 | 1.76T | 815B | 1.51T | 0 | 6.84T |
| 2026-08-12 | 1.72T | 871B | 1.58T | 0 | 7.06T |
| 2026-08-13 | 1.63T | 790B | 1.50T | 0 | 7.21T |
| 2026-08-14 | 1.62T | 746B | 1.42T | 0 | 7.70T |
| 2026-08-15 | 1.41T | 652B | 1.12T | 0 | 6.39T |
| 2026-08-16 | 1.46T | 648B | 1.06T | 0 | 7.21T |
| 2026-08-17 | 1.69T | 1.05T | 1.52T | 0 | 9.06T |
| 2026-08-18 | 1.74T | 1.04T | 1.63T | 0 | 8.29T |
| 2026-08-19 | 1.68T | 807B | 1.22T | 0 | 8.41T |
| 2026-08-20 | 1.83T | 656B | 1.25T | 0 | 8.39T |
| 2026-08-21 | 1.85T | 555B | 1.12T | 1.99T | 8.37T |
| 2026-08-22 | 1.39T | 417B | 754B | 4.54T | 6.97T |
| 2026-08-23 | 1.39T | 386B | 707B | 5.02T | 7.64T |
| 2026-08-24 | 1.71T | 613B | 1.07T | 5.93T | 9.03T |
| 2026-08-25 | 1.77T | 724B | 1.03T | 5.75T | 9.50T |
| 2026-08-26 | 2.41T | 753B | 1.04T | 4.00T | 10.2T |
| 2026-08-27 | 2.00T | 1.45T | 1.11T | 0 | 11.2T |
| 2026-08-28 | 1.61T | 1.70T | 963B | 0 | 10.8T |
| 2026-08-29 | 1.37T | 1.12T | 699B | 0 | 10.0T |
| 2026-08-30 | 1.44T | 1.43T | 748B | 0 | 9.70T |
| 2026-08-31 | 1.59T | 1.30T | 825B | 0 | 10.6T |
| 2026-09-01 | 1.63T | 1.76T | 515B | 0 | 12.3T |
| 2026-09-02 | 1.73T | 2.84T | 730B | 0 | 13.2T |
| 2026-09-03 | 1.91T | 1.44T | 768B | 0 | 14.4T |
| 2026-09-04 | 2.24T | 1.39T | 510B | 0 | 14.1T |
| 2026-09-05 | 1.73T | 2.01T | 343B | 0 | 11.3T |
| 2026-09-06 | 1.52T | 2.19T | 301B | 0 | 10.3T |
| 2026-09-07 | 1.58T | 2.56T | 495B | 0 | 14.4T |
| # | Model | Tokens |
|---|---|---|
| 1 | DeepSeek: DeepSeek V4 Flash 0731 (batch)deepseek 11.8% share · $0.14/M in · 1049K ctx | 50.3T |
| 2 | OpenAI: GPT-5.6 Luna (batch)openai 8.0% share · $0.10/M in · 1050K ctx | 34.1T |
| 3 | Tencent: Hy3tencent 7.3% share · $0.13/M in · 262K ctx | 30.9T |
| 4 | Ox Alphastealth 6.4% share | 27.2T |
| 5 | Xiaomi: MiMo-V2.5xiaomi 6.2% share · $0.14/M in · 1050K ctx | 26.6T |
| 6 | DeepSeek: DeepSeek V4 Flash 0423deepseek 5.2% share · $0.09/M in · 1049K ctx | 22.0T |
| 7 | Tencent: Hy4 previewtencent 5.0% share · $0.83/M in · 1049K ctx | 21.1T |
| 8 | Z.ai: GLM 5.3 Flash (batch)z ai 4.8% share · $0.15/M in · 1049K ctx | 20.4T |
| 9 | Nemotron 3 Ultra 550b A55b 20260604:Freenvidia 3.9% share | 16.8T |
| 10 | Z.ai: GLM 5.2z ai 3.3% share · $0.97/M in · 1049K ctx | 13.9T |
| 11 | Google: Gemini 3.7 Flash (batch)google 2.0% share · $0.38/M in · 1049K ctx | 8.35T |
| 12 | Minimax M3 20260531:Freeminimax 1.9% share | 8.17T |
| 13 | DeepSeek: DeepSeek V4 Pro 0423deepseek 1.9% share · $0.96/M in · 1049K ctx | 8.08T |
| 14 | Claude Opus 5 (batch)anthropic 1.8% share · $2.50/M in · 1000K ctx | 7.85T |
| 15 | MoonshotAI: Kimi K3 (batch)moonshotai 1.6% share · $3.00/M in · 1049K ctx | 6.70T |
| 16 | MiniMax: MiniMax M3 (batch)minimax 1.6% share · $0.30/M in · 524K ctx | 6.63T |
| 17 | Laguna S 2.1 20260720:Freepoolside 1.5% share | 6.34T |
| 18 | OpenAI: GPT-5.6 Sol (batch)openai 1.4% share · $1.00/M in · 1050K ctx | 6.00T |
| 19 | Z.ai: GLM 5.3z ai 1.3% share · $1.40/M in · 1311K ctx | 5.38T |
| 20 | Anthropic: Claude Sonnet 5 (batch)anthropic 1.2% share · $1.00/M in · 1000K ctx | 4.98T |
All models outside the top 50 accounted for 24.5T tokens in this window.
Artificial Analysis composite score, all evaluated models
| 1 | Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 53.4 |
| 2 | Qwen3.8 Max | 53.4 |
| 3 | GPT-6 Astra (max) | 52.8 |
| 4 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | 50.7 |
| 5 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | 49.7 |
| 6 | GPT-5.6 Sol (max) | 47.1 |
| 7 | GLM-5.3 (max) | 44.9 |
| 8 | Grok 4.6 (high) | 44.4 |
| 9 | Kimi K3 (max) | 43.8 |
| 10 | GPT-5.6 Terra (max) | 42.3 |
OpenRouter's own eval — graduate-level science questions
| 1 | Google: Gemini 3.1 Pro Preview | 94.4% |
| 2 | OpenAI: GPT-6 Astra | 94.4% |
| 3 | Google: Gemini 3.7 Flash | 94.3% |
| 4 | OpenAI: GPT-5.5 | 93.8% |
| 5 | OpenAI: GPT-5.6 Sol Pro | 93.8% |
| 6 | SpaceXAI: Grok 4.6 | 93.3% |
| 7 | Google: Gemini 3.6 Flash | 92.8% |
| 8 | Google: Gemini 3.5 Flash | 92.8% |
| 9 | OpenAI: GPT-5.6 Sol | 91.9% |
| 10 | MoonshotAI: Kimi K3 | 91.5% |
Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).
Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings).
The Mr. Bright Role Index
Which model should do which job?
Leaderboards rank models against each other. Businesses hire them for specific jobs. We score every benchmarked model against the marketing roles we actually staff — weighting the signals that predict performance in that seat, not overall cleverness. Weights are published below each role so you can check the maths.
Software Engineer
anthropic · $2.50/M
- 2. OpenAI: GPT-5.6 Sol (batch)91.4
- 3. MoonshotAI: Kimi K3 (batch)86.9
Code correctness dominates; tool use matters because the job is multi-step.
Web Designer & CRO
anthropic · $2.50/M
- 2. OpenAI: GPT-5.6 Sol (batch)90.3
- 3. Z.ai: GLM 5.386.4
Ships front-end code and reasons about experiments, so coding and reasoning share the load.
Paid Social Buyer
z ai · $0.150/M
- 2. OpenAI: GPT-5.6 Luna (batch)69.8
- 3. Claude Opus 5 (batch)65.2
Runs long tool-driven loops over ad platforms every day, so agentic ability and cost per run lead.
Search & PPC Manager
z ai · $0.150/M
- 2. OpenAI: GPT-5.6 Luna (batch)69.2
- 3. Claude Opus 5 (batch)65.2
Same daily tool loop as paid social, with more numeric reasoning over search terms.
SEO & AI Search
openai · $1.00/M
- 2. Google: Gemini 3.7 Flash (batch)78.8
- 3. MoonshotAI: Kimi K3 (batch)78.8
Reads large corpora and must be factually tight, so context length and hard-reasoning accuracy carry weight.
Lifecycle Marketer
openai · $0.100/M
- 2. Z.ai: GLM 5.3 Flash (batch)69.1
- 3. Claude Opus 5 (batch)60.2
High message volume makes cost per token decisive alongside writing quality.
Creative Director
openai · $0.100/M
- 2. Z.ai: GLM 5.3 Flash (batch)67.6
- 3. Z.ai: GLM 5.362.9
Generates constantly and holds long brand context, so throughput economics and context matter as much as raw capability.
Community Manager
openai · $0.100/M
- 2. Z.ai: GLM 5.3 Flash (batch)67
- 3. OpenAI: GPT-5.6 Sol (batch)62.5
Answers the public all day: cheap, accurate and steady beats brilliant and expensive.
Manager & Orchestrator
anthropic · $2.50/M
- 2. Z.ai: GLM 5.389.3
- 3. OpenAI: GPT-5.6 Sol (batch)83.7
Reviews other agents' work and coordinates them, which is the hardest reasoning and longest context job on the team.
How the index is calculated
Every signal is min-max normalised to 0–100 across the benchmarked models in the current window, then combined using the weights shown on each card. Cost efficiency is the inverse of prompt price; context is log-scaled so a 1M-token window doesn't swamp every other signal. A model is only ranked for a role when at least half that role's weight is backed by published data, and scores are re-based on the weight actually available — so a model isn't penalised for a benchmark nobody has run on it. Models without an intelligence score are excluded entirely.
This index is our own analysis built on public data. It is a starting point for choosing a brain, not a guarantee of performance on your workload — the right model for your account is the one that wins your own evaluation.
Usage data: OpenRouter public rankings dataset. Token totals are prompt + completion across all public traffic. Mr. Bright AI is not affiliated with OpenRouter.
Usage, not hype
Rankings are ordered by real tokens processed across OpenRouter's public traffic — prompt plus completion — not by marketing claims or vibes.
Benchmarks alongside
Each model carries its Artificial Analysis Intelligence and Coding indexes plus OpenRouter's own GPQA Diamond eval, so you can weigh capability against adoption.
Momentum tracked
The growth column compares the second half of the window against the first, so you can spot models being adopted — or abandoned — before the rank flips.
Refreshed daily
The usage dataset rebuilds every day; benchmark indexes update as sources publish new runs. No stale screenshots.
We run on the models at the top of this board.
Every Mr. Bright AI team is model-agnostic — we route each task to whichever model is winning at it right now, and we re-check this data continuously.
See our full stack