What Drives Pricing Differences
Model size and capability tier, input vs. output token pricing (output is often priced higher), and whether a model is open-weight (self-hosted, no per-token fee) or closed (managed API with per-token billing) all factor into the price you actually pay.
General Pricing Tiers Across Providers
Budget-tier models (GPT-4o mini, Gemini 2.0 Flash, Claude Haiku) generally cost a fraction of flagship-tier models (GPT-5, Claude Opus 4.8, Gemini 2.5 Pro) per token — see our Pricing directory for current exact rates by platform.
Estimating Your Actual Cost
See our LLM Cost Calculator for estimating real cost based on your expected usage volume, since headline per-token rates alone don't tell the full story.
Related Pages
Frequently Asked
Why is output token pricing often higher than input?
Generating text is generally more computationally intensive than processing input, which is reflected in most providers' pricing structures.
Is open-weight always cheaper than a closed API?
Often yes at scale, though you need to factor in your own infrastructure cost if self-hosting rather than using a managed API.
Where can I see current exact pricing by provider?
See our Pricing directory for current tier breakdowns.
How do I estimate my actual expected cost?
See our LLM Cost Calculator.