Token Cost Calculator

Token Cost Calculator

Compare prices across OpenAI, Claude, Gemini, DeepSeek, and Grok. Real tokenizer for OpenAI, accurate estimates for the rest.

Last updated: 2026-09-26

1,000
2,000
400
1,000 × 2,000in + 400out = 2,000,000 input,400,000 output tokens

Estimated cost across 18 models

google Cheapest

Gemini 2.5 Flash-Lite

$0.360/ month
$4.32 / year
openai

GPT-6 luna

$0.400/ month
$4.80 / year
deepseek

DeepSeek V4.1-Flash

$1.08/ month
$12.96 / year
Peak hours
google

Gemini 3.1 Flash-Lite

$1.10/ month
$13.20 / year
google

Gemini 2.5 Flash

$2.70/ month
$32.40 / year
Long context ≥200,000
google

Gemini 3.8 Flash

$3.00/ month
$36.00 / year
anthropic

Claude Haiku 4.5

$4.00/ month
$48.00 / year
deepseek

DeepSeek V4-Pro

$4.22/ month
$50.69 / year
Peak hours
grok

Grok 4.7

$6.40/ month
$76.80 / year
openai

GPT-6 sol

$8.00/ month
$96.00 / year
anthropic

Sonnet 5

$8.00/ month
$96.00 / year
google

Gemini 3.1 Pro Preview

$8.80/ month
$106 / year
google

Gemini 2.5 Pro

$11.00/ month
$132 / year
Long context ≥200,000
openai

GPT-5.6 sol

$16.00/ month
$192 / year
anthropic

Opus 5.5

$16.00/ month
$192 / year
openai

GPT-6 astra

$40.00/ month
$480 / year
anthropic

Fable 5.1

$40.00/ month
$480 / year
openai

GPT-5.6 cyber

$55.00/ month
$660 / year

Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json

Quick takeaways
  • 1Output tokens usually cost 4–8× more than input — keeping replies short is the highest-leverage cost saver.
  • 2Anthropic prompt caching drops repeat-input cost by ~90%; OpenAI cached input drops it by ~90% on GPT-5.
  • 3DeepSeek off-peak hours (16:30–00:30 UTC) are 10× cheaper than peak — schedule batch jobs accordingly.
  • 4Gemini 2.5 Pro doubles its input price above 200K context — large-context calls cost roughly twice as much.
Token Cost Calculator

How to bring the bill down

  1. 1Use a smaller model for routing / classification — Haiku 4.5 / GPT-5 nano cost roughly 5–10× less than the flagship.
  2. 2Cache repeat prefixes (system prompts, tool schemas) — Anthropic cached input drops the rate by ~90%.
  3. 3Set max_tokens explicitly — output is the expensive side; capping it prevents runaway replies.
  4. 4Batch non-urgent jobs into DeepSeek off-peak hours (16:30–00:30 UTC) for ~10× cheaper inference.
  5. 5Trim long context — Gemini 2.5 Pro doubles its input rate above 200K; chunk documents that don't need the rest.
Token Cost Calculator

calculator.faqTitle

It depends entirely on the model. GPT-5 nano is around $0.05 input / $0.40 output per 1M tokens; Claude Opus 4.5 is $5 / $25; Gemini 2.5 Flash is $0.30 / $2.50. Use the calculator above to compare every major model at your exact usage.