Every vendor deck names a different best AI model, and none of them shows you usage. Usage is countable, and one of the few places it is published is OpenRouter, a marketplace that routes requests to more than 500 models and publishes what flows through it. Measured by tokens processed through OpenRouter in late September 2026, the most used AI models are cheap flash models from Chinese labs, led by DeepSeek V4.1 Flash and GLM 5.3 Flash. The flagships you read about sit far down the table, and two of the top ten are free. This post walks the ranking as fetched on 28 September 2026, the prices that explain it, and what the numbers do and do not measure.
How the ranking was built
The numbers below are OpenRouter’s model rankings, fetched on 28 September 2026. The methodology is published on the page: “Each model is ranked by the number of tokens it processed through the OpenRouter API, counting both prompt and completion tokens.” The main leaderboard covers the trailing seven days, ending with the most recent complete day, and variants of the same model, such as a free variant, are ranked separately. Every token figure and percentage in this post comes from that page and nowhere else.
The scale is large enough to mean something: OpenRouter’s own homepage counts more than 500 trillion tokens a month across over 10 million users and more than 80 providers. Simon Willison’s September 2026 note on OpenRouter quotes the platform’s own selling point, one endpoint that handles fallbacks and picks the most cost-effective backend for each request, which tells you who shops there: teams that treat models as interchangeable and route on price.
That is also the built-in bias. OpenRouter states it plainly: “These rankings show traffic routed through OpenRouter.” A company on a direct OpenAI, Anthropic or Google contract never appears in this data, so the board over-represents price-sensitive, multi-model workloads and under-represents single-vendor enterprises. It measures neither users nor spend, by its own description. Read it as the spot market of AI, not the whole market.
Chinese flash models lead, and first place is eighteen days old
First place is eighteen days old. DeepSeek released V4.1 Flash on 10 September 2026 as a 552B-parameter mixture-of-experts model with, per its own announcement, 8B active parameters for input and 16B for output. The same announcement set V4-Pro on a phase-out, its requests routed to the new model until a V4.1-Pro launches: “Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime.” Two weeks later it processed 19.6 trillion tokens in seven days on OpenRouter, up 24 percent week over week, and DeepSeek holds two more of the ten slots with the older V4 Flash builds it is migrating away from.
Second place is the same play from Z.ai. Its GLM 5.3 Flash page describes a 320B-parameter model with 18B activated, and calls it “the first open-source frontier model to combine sparse and linear attention”. It carried 16.3 trillion tokens in the week, and it leads the monthly table at 57.5 trillion. Because the weights are open, OpenRouter lists 31 providers serving it, with a floor price of $0.045 per million input tokens, less than a third of Z.ai’s own first-party rate. That is what an open-weight model does to serving prices: the model becomes a file, and hosts compete the margin away.
Add the ten rows together and the board carried 93.2 trillion tokens in seven days, with the two Chinese flash leaders alone at 35.9 trillion, more than a third of it. Six of the ten slots belong to models from Chinese labs: three from DeepSeek, one each from Z.ai, Tencent and Xiaomi.
Third place is anonymous, and two of the ten are free
The strangest row is the third one. Space Bunny Alpha is listed under the author “stealth”: an anonymous model with a million-token context window, free to use, that processed 13.9 trillion tokens in its first week on the board. OpenRouter lists it as a preview from a provider that has chosen to stay anonymous, and the volume shows how much traffic a price of zero attracts regardless of the logo on it. NVIDIA’s Nemotron 3 Ultra free variant sits at number seven on the same logic.
The trap: a free tier is a promotion, not a price. Volume that exists because something costs nothing tells you what developers will try, not what a business should run, and the free rows on this board are the clearest case of the metric measuring generosity instead of adoption.
Where the famous names sit
The Western presence in the weekly top ten is thin: OpenAI’s budget line, with GPT-5.6 Luna at number five on 8.53 trillion tokens and GPT-6 Luna entering at number ten with 2.86 trillion within days of release, plus the free NVIDIA variant above. By share of requests, OpenRouter’s author table for the week beginning 21 September 2026 puts DeepSeek first at 24.4 percent, Google second at 20.5 percent and OpenAI third at 17.9 percent.
Google’s position is the instructive one: a fifth of all requests, yet no Google model in the token top ten. Requests and tokens are different meters, and a model that answers many short calls can lead one board and miss the other. Anthropic is the other lesson, at 2.7 percent of requests even as its brand-new Claude Opus 5.5 entered OpenRouter’s new-and-trending table with 1.08 trillion tokens in its first week. Anthropic’s models are priced at the top of the market, and a marketplace whose users route on price is exactly where premium models look smallest. A famous name missing from this board is usually being bought through a different door, on direct contracts this data cannot see.
The price war underneath the table
The ranking makes sense the moment you put prices beside it. On 22 September 2026, Simon Willison recorded a day in which Anthropic released Claude Opus 5.5 and OpenAI answered with GPT-6 Sol and GPT-6 Luna about an hour later, the new GPT-6 models at half the price of their GPT-5.6 equivalents and Opus taking its first price cut in five generations, 20 percent down to 4 and 20 dollars per million tokens. His verdict: “It’s hard to overstate how competitive this pricing is.”
The floor of that war is where the volume lives. OpenAI’s pricing page lists GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output. DeepSeek’s price sheet lists V4.1 Flash at $0.30 input and $1.20 output at peak, half that off-peak, with cache hits from $0.003. Z.ai charges $0.15 and $0.50 for GLM 5.3 Flash. Set those against the top of Willison’s same price table, where Claude Fable 5.1 and GPT-6 Astra both list at 10 and 50 dollars per million: the spread between the most expensive frontier model and the cheapest current one is now a factor of 100, on both input and output. A two-order-of-magnitude price range across working models is why token volume pools at the bottom of it.
What the token numbers hide
OpenRouter publishes its own caveats, and the central one deserves quoting: models differ in verbosity and tokenization, so “a higher token total shows how much a model is used, not which model is best for a task”. Nothing on this board measures accuracy or reasoning ability, and a chatty model literally counts more than a terse one on the same work.
Three more distortions matter before you quote any of this. The board is OpenRouter-shaped: direct enterprise traffic to the big labs is invisible, which flatters the price-driven tail. Free tiers inflate their rows, as above. And the table is volatile: the two top models are weeks old, the change column swings 20 percent week over week, and the monthly view already tells a different story, with GLM 5.3 Flash first and the top weekly model only fifth because it is too new to have a full month. A ranking this liquid is a snapshot, which is why this post carries its fetch date in every figure.
How to actually use this ranking
For a buyer, the table answers one question well: what does the market’s working tier look like when someone else is paying attention to price? Three uses survive the caveats:
- Calibrate the default. The most used models are not the household names, and the household names are not overpriced by accident: they sell capability, the flash tier sells throughput. If your workloads never distinguish the two, you are paying the capability premium for nothing.
- Split your traffic deliberately. The pattern the board rewards is model routing: routine, high-volume steps on a flash-class model at cents per million tokens, the hard steps on a frontier model. The 100x spread is the budget you recover by doing this.
- Keep the exit open. This week’s number one is eighteen days old, and the list will reorder before your next quarterly review. A multi-provider setup behind your own gateway turns model choice into configuration, and if the cheap rows tempt you toward the Chinese labs, price is the smallest of the three questions to clear first; the other two are in Can you use a Chinese AI model?
The same method drives our framework rankings: registry data over vendor claims, with the distortions named. The Python edition is in the 6 most used Python AI agent frameworks.
The takeaway
The most used AI models in late September 2026 are the cheapest capable ones: DeepSeek V4.1 Flash and GLM 5.3 Flash lead OpenRouter’s board, OpenAI shows up with its budget Luna line, and the premium flagships are bought elsewhere, on direct contracts the data cannot see. Usage follows price wherever price is allowed to matter, and the spread between the frontier and the flash tier now runs to a factor of 100. Treat the board as the spot market it is: check it on the day you decide, route on it deliberately, and never mistake token volume for a verdict on quality. Ranked usage narrows a shortlist rather than picking for you, the same conclusion the registry forced on us in the TypeScript framework ranking.
Sources
- LLM rankings, OpenRouter: tokens processed, author market share and methodology, fetched 28 September 2026.
- DeepSeek-V4.1-Flash release, DeepSeek news: the 10 September 2026 launch and the V4-Pro retirement.
- GLM-5.3-Flash, Z.AI developer documentation: the model’s architecture and open-source claim.
- Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war, Simon Willison: the 22 September releases and the full price table.
- OpenAI API pricing: the GPT-6 Luna rates.
- Models and pricing, DeepSeek API docs: the V4.1 Flash rates and the peak and off-peak split.
- Pricing, Z.AI developer documentation: the first-party GLM-5.3-Flash rates.
- OpenRouter: the monthly token, user and provider counts.
- GLM 5.3 Flash on OpenRouter: the provider count and the floor price.
- Space Bunny Alpha on OpenRouter: the anonymous preview listing.
- So you want to use OpenRouter?, Simon Willison: how the marketplace routes requests across providers.