Token Cost Calculator · Free

See what your LLM tokens actually cost — across every major model.

Compare prices across OpenAI, Claude, Gemini, DeepSeek, and Grok in one place. Real tokenizer for OpenAI models, accurate estimates for the rest. Updated regularly.

Trusted by engineers comparing LLM costs every day

100k input · 40k output · 1k req/mo
6 of 18 models shown
Updated 2026-09-26
Gemini 2.5 Flash-Lite Cheapest
$26.00/ month
GPT-6 luna
$30.00/ month
DeepSeek V4.1-Flash
$78.00/ month
Gemini 3.1 Flash-Lite
$85.00/ month
Gemini 2.5 Flash
$210.00/ month
Gemini 3.8 Flash
$225.00/ month
Sorted by cheapest first · 18 models · 5 providersSource: provider pricing pages
1,000
2,000
400
1,000 × 2,000in + 400out = 2,000,000 input,400,000 output tokens

Estimated cost across 18 models

google Cheapest

Gemini 2.5 Flash-Lite

$0.360/ month
$4.32 / year
openai

GPT-6 luna

$0.400/ month
$4.80 / year
deepseek

DeepSeek V4.1-Flash

$1.08/ month
$12.96 / year
Peak hours
google

Gemini 3.1 Flash-Lite

$1.10/ month
$13.20 / year
google

Gemini 2.5 Flash

$2.70/ month
$32.40 / year
Long context ≥200,000
google

Gemini 3.8 Flash

$3.00/ month
$36.00 / year
anthropic

Claude Haiku 4.5

$4.00/ month
$48.00 / year
deepseek

DeepSeek V4-Pro

$4.22/ month
$50.69 / year
Peak hours
grok

Grok 4.7

$6.40/ month
$76.80 / year
openai

GPT-6 sol

$8.00/ month
$96.00 / year
anthropic

Sonnet 5

$8.00/ month
$96.00 / year
google

Gemini 3.1 Pro Preview

$8.80/ month
$106 / year
google

Gemini 2.5 Pro

$11.00/ month
$132 / year
Long context ≥200,000
openai

GPT-5.6 sol

$16.00/ month
$192 / year
anthropic

Opus 5.5

$16.00/ month
$192 / year
openai

GPT-6 astra

$40.00/ month
$480 / year
anthropic

Fable 5.1

$40.00/ month
$480 / year
openai

GPT-5.6 cyber

$55.00/ month
$660 / year

Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json

Quick takeaways
  • 1Output tokens usually cost 4–8× more than input — keeping replies short is the highest-leverage cost saver.
  • 2Anthropic prompt caching drops repeat-input cost by ~90%; OpenAI cached input drops it by ~90% on GPT-5.
  • 3DeepSeek off-peak hours (16:30–00:30 UTC) are 10× cheaper than peak — schedule batch jobs accordingly.
  • 4Gemini 2.5 Pro doubles its input price above 200K context — large-context calls cost roughly twice as much.
Token Cost Calculator

How to bring the bill down

  1. 1Use a smaller model for routing / classification — Haiku 4.5 / GPT-5 nano cost roughly 5–10× less than the flagship.
  2. 2Cache repeat prefixes (system prompts, tool schemas) — Anthropic cached input drops the rate by ~90%.
  3. 3Set max_tokens explicitly — output is the expensive side; capping it prevents runaway replies.
  4. 4Batch non-urgent jobs into DeepSeek off-peak hours (16:30–00:30 UTC) for ~10× cheaper inference.
  5. 5Trim long context — Gemini 2.5 Pro doubles its input rate above 200K; chunk documents that don't need the rest.
frequently asked

Questions about the calculator

It depends entirely on the model. GPT-5 nano is around $0.05 input / $0.40 output per 1M tokens; Claude Opus 4.5 is $5 / $25; Gemini 2.5 Flash is $0.30 / $2.50. Use the calculator above to compare every major model at your exact usage.

Ready to compare your LLM bill?

Paste your prompt, set your usage, see the cost across every model — in seconds. No signup.