Market preview
DeepSeek: DeepSeek Pro Latest+27%DeepSeek V4 Flash Latest+16%DeepSeek: DeepSeek V4 Flash 0731+16%@cf/ai4bharat/indictrans2-en-indic-1B$0.342 / 1MV100$0.02/hr
Five identical model cartridges sold through different windows, with published Mistral Small 3 input prices ranging from $0.05 to $0.90 per million tokens.
Conceptual editorial image · not report evidence
StackTicker structured report methodology
StackTicker structured report methodology

See where this value came from and when it was checked.

Source type
official
Source retrieved
Sep 1, 2026 · 16:09 UTC
Effective date
Not stated
Verification
Source verified
View source ↗

A million input tokens of Mistral Small 3 costs $0.05 from one provider and $0.90 from another. It is the same model. The weights are public, the architecture is published, and the weights are the same weights.

Open weights were supposed to make inference a commodity. They made the model a commodity. The price never followed.

Open weights made the model a commodity. They did not make the price one - the same open model is sold at up to 18x depending on who you happen to ask.

Same open model, cheapest and dearest provider

Lowest published input price per million tokens, per provider. Models offered by at least three providers we track.

ModelProvidersCheapestMost expensiveMultiple
Mistral: Mistral Small 33$0.05/M · OpenRouter$0.90/M · Fireworks AI18.0x
Google: Gemma 3 27B3$0.08/M · OpenRouter$0.90/M · Fireworks AI11.2x
Qwen: Qwen3 32B3$0.08/M · DeepInfra$0.90/M · Fireworks AI11.2x
Google: Gemma 4 31B3$0.09/M · OpenRouter$0.90/M · Fireworks AI10.0x
Qwen: Qwen3 Next 80B A3B Instruct3$0.09/M · DeepInfra$0.90/M · Fireworks AI10.0x
DeepSeek: DeepSeek V4 Flash 04234$0.05/M · OpenRouter$0.44/M · DeepSeek9.0x

Why the gap survives

Three reasons, and only the first is about serving costs. Providers genuinely differ in hardware, batching and utilisation, and a provider running a model badly does pay more per token.

The second is that the list price and the price people pay are not the same number. Committed spend, reserved capacity and enterprise agreements are negotiated privately, so a published rate is partly a starting position. We publish the rate a new customer would meet, because it is the only one that exists in public.

The third is the one that matters: almost nobody checks. When a model is available from a dozen providers, comparing them is a morning's work spread across a dozen pricing pages that use different units and different token accounting. Most teams pick a provider once, early, for reasons that were true at the time, and then stop looking.

Our view

A price difference that survives only because comparison is tedious is not a market outcome, it is a friction rent. It disappears the moment the comparison becomes easy - which is the entire reason this site exists, and we would rather say that plainly than pretend to be neutral about it.

None of this makes the expensive provider dishonest. Latency, throughput, uptime and data policy differ, and some of the premium buys something real. But those are claims a buyer should get to weigh against a number, and the number has been the hard part.

The check worth making

Take the open model you use most. Find what you pay per million tokens. Then find the cheapest provider serving the same weights. If the ratio is above two, you are paying for something - and you should be able to say what.

Run a modelCompare providers by price

The same model can cost several times more depending on who serves it. See every provider we track for a model, with the price each one publishes and the page it came from.

Open the comparison →
Run it yourselfCheapest GPUs by the hour

Renting the hardware keeps the weights, the prompts and the logs on infrastructure you choose. See what an hour of each accelerator costs across every host we track.

Open the GPU market →