Analysis
One model, an 18x spread in published API prices
Mistral Small 3 was listed from $0.05 to $0.90 per million input tokens in our snapshot. The model is the same; the service around it is not.

StackTicker structured report methodology
See where this value came from and when it was checked.
- Source type
- official
- Source retrieved
- Sep 1, 2026 · 16:09 UTC
- Effective date
- Not stated
- Verification
- Source verified
A million input tokens of Mistral Small 3 costs $0.05 from one provider and $0.90 from another. It is the same model. The weights are public, the architecture is published, and the weights are the same weights.
Open weights were supposed to make inference a commodity. They made the model a commodity. The price never followed.
Open weights made the model a commodity. They did not make the price one - the same open model is sold at up to 18x depending on who you happen to ask.
Same open model, cheapest and dearest provider
Lowest published input price per million tokens, per provider. Models offered by at least three providers we track.
| Model | Providers | Cheapest | Most expensive | Multiple |
|---|---|---|---|---|
| Mistral: Mistral Small 3 | 3 | $0.05/M · OpenRouter | $0.90/M · Fireworks AI | 18.0x |
| Google: Gemma 3 27B | 3 | $0.08/M · OpenRouter | $0.90/M · Fireworks AI | 11.2x |
| Qwen: Qwen3 32B | 3 | $0.08/M · DeepInfra | $0.90/M · Fireworks AI | 11.2x |
| Google: Gemma 4 31B | 3 | $0.09/M · OpenRouter | $0.90/M · Fireworks AI | 10.0x |
| Qwen: Qwen3 Next 80B A3B Instruct | 3 | $0.09/M · DeepInfra | $0.90/M · Fireworks AI | 10.0x |
| DeepSeek: DeepSeek V4 Flash 0423 | 4 | $0.05/M · OpenRouter | $0.44/M · DeepSeek | 9.0x |
Why the gap survives
Three reasons, and only the first is about serving costs. Providers genuinely differ in hardware, batching and utilisation, and a provider running a model badly does pay more per token.
The second is that the list price and the price people pay are not the same number. Committed spend, reserved capacity and enterprise agreements are negotiated privately, so a published rate is partly a starting position. We publish the rate a new customer would meet, because it is the only one that exists in public.
The third is the one that matters: almost nobody checks. When a model is available from a dozen providers, comparing them is a morning's work spread across a dozen pricing pages that use different units and different token accounting. Most teams pick a provider once, early, for reasons that were true at the time, and then stop looking.
Our view
A price difference that survives only because comparison is tedious is not a market outcome, it is a friction rent. It disappears the moment the comparison becomes easy - which is the entire reason this site exists, and we would rather say that plainly than pretend to be neutral about it.
None of this makes the expensive provider dishonest. Latency, throughput, uptime and data policy differ, and some of the premium buys something real. But those are claims a buyer should get to weigh against a number, and the number has been the hard part.
The check worth making
Take the open model you use most. Find what you pay per million tokens. Then find the cheapest provider serving the same weights. If the ratio is above two, you are paying for something - and you should be able to say what.
The same model can cost several times more depending on who serves it. See every provider we track for a model, with the price each one publishes and the page it came from.
Open the comparison →Run it yourselfCheapest GPUs by the hourRenting the hardware keeps the weights, the prompts and the logs on infrastructure you choose. See what an hour of each accelerator costs across every host we track.
Open the GPU market →