01 · Sources
Where a figure came from
Every published figure carries the kind of source it came from. These are the five names the site uses, and they appear in the evidence panel beside the value.
Read from the provider’s own machine-readable endpoint.
The provider’s own pricing page, product page or docs.
A reseller’s listing for somebody else’s model. Real, and not the model owner’s own rate.
A published evaluation, attributed to whoever ran it.
Reported by a third party. Used when nothing better exists, and labelled.
02 · Time
Three dates that mean different things
Prices are stored append-only: an unchanged rate writes no new row. That makes one date insufficient, so the site keeps three and never uses one word for all of them.
When this price first appeared, or last moved.
When we last read the source and found it unchanged. This is the only date behind a “verified today” badge.
When the document behind the figure was fetched.
03 · Certainty
Not all claims are equally certain
The observation passed validation against its expected unit, scope and entity.
More than one independent source supports the same underlying fact.
StackTicker derived the value from disclosed inputs, which are shown.
No source publishes this value. It is shown as absent, never as zero.
04 · Pricing
Normalised without losing provenance
A provider offering owns the pricing structure, not the model in isolation. The raw value, currency, unit, region, tier and both timestamps stay available beside every normalised comparison.
Model → Provider offering → Pricing observation → Source
05 · Stability
Repeated evidence, not sentiment
Capability stability uses fixed, difficult canaries that depend on supplied evidence and expose long-horizon reasoning failures. A score needs enough observations, a time window, a method and a confidence before it publishes.
What can change
Hypothesis tracking, causal consistency, contradictions, instruction adherence, coding consistency and tool reliability.
What you see
Stable, Watch, Regression, Recovering — or not enough observations yet, which is stated rather than hidden behind a number.
06 · Rankings
Intent before leaderboard
A ranking names the goal it serves, discloses its metric and its update time, and does not declare a winner where the difference is unsupported or the things are not comparable.
07 · Analysis
Fact separated from inference
- Observed fact: directly supported by evidence.
- Strong inference: likely, given repeated behaviour and incentives.
- Plausible hypothesis: useful, not yet confirmed.
- Speculation: marked as such, and never presented as data.
