LLM API pricing · live reference
Stop guessing from pricing pages. Plug in your real workload and see every major model — Claude, GPT, Gemini, DeepSeek — ranked by what you'll actually pay.
Share of input served from prompt cache. Only applies to models that offer cached pricing.
DeepSeek V4 Flash
$26.29
/ month
Monthly estimate = per-request cost × requests per day × 30. Cached input billed at each model's cache rate where offered. Figures are estimates — verify against official pricing before committing spend.
The matchups people search for most.
| Model | Provider | Input /1M | Output /1M | Cached /1M | Context |
|---|---|---|---|---|---|
Claude Opus 4.8Frontier | Anthropic | $5.00 | $25.00 | $0.500 | 200K |
Claude Sonnet 4.6Balanced | Anthropic | $3.00 | $15.00 | $0.300 | 200K |
Claude Haiku 4.5Fast | Anthropic | $1.00 | $5.00 | $0.100 | 200K |
GPT-5.5Frontier | OpenAI | $5.00 | $30.00 | — | 400K |
GPT-5.4Balanced | OpenAI | $2.50 | $15.00 | — | 400K |
OpenAI o3Frontier | OpenAI | $2.00 | $8.00 | — | 200K |
GPT-4.1Balanced | OpenAI | $2.00 | $8.00 | — | 1M |
GPT-4.1 nanoFast | OpenAI | $0.100 | $0.400 | — | 1M |
Gemini 3.1 ProFrontier | $2.00 | $12.00 | — | 1M | |
Gemini 2.5 ProBalanced | $1.25 | $10.00 | $0.310 | 1M | |
Gemini 2.5 FlashFast | $0.300 | $2.50 | $0.075 | 1M | |
DeepSeek V4 FlashOpen | DeepSeek | $0.140 | $0.280 | $0.014 | 128K |
DeepSeek V4 ProOpen | DeepSeek | $1.74 | $3.48 | — | 128K |
Grok 4.3Balanced | xAI | $1.25 | $2.50 | — | 256K |
Grok 4.1 FastFast | xAI | $0.200 | $0.500 | — | 256K |
Verified 2026-06-21· Prices in USD per million tokens. Always confirm on the provider's page before committing spend.
Need savings this month?
If your bill is already real, the calculator is only the first pass. Send rough usage numbers or a provider export and get a short written report: likely waste, cheaper-model swaps, cache opportunities, and the next experiment to run. If the cost comes from agent loops, RAG over-retrieval, model routing drift, retry storms, cache misses, or coding-agent tool calls, choose the $299 Cost Leak Review path.
24h
Reply target after written scope
No keys
Usage summaries are enough
$99/$299/$1k
Paid only after scope acceptance
Agent cost leak and emergency path
The $299 path is for one bounded agent cost leak. The $1,000 emergency sprint is for runaway LLM bills where agent loops, RAG fan-out, retry storms, cache misses, routing drift, or launch-pricing risk need a fast containment plan.
Best fit: teams spending at least a few hundred dollars per month on OpenAI, Anthropic, Gemini, DeepSeek, or router APIs. This is optimization guidance, not provider billing support.
TokenMeter Pro · founder waitlist
Pro will turn the calculator into a live AI spend dashboard: connect read-only provider usage keys, see monthly spend across Claude and OpenAI first, get budget alerts, and find cheaper model swaps before the next bill lands.
Waitlist signups get the founder rate; no payment until the product is ready and you choose to subscribe.