Token Cost Calculator

OpenRouter Pricing Calculator

See what your tokens actually cost after the OpenRouter markup — and when going direct to the provider is cheaper.

Last updated: 2026-09-26

1,000
2,000
400
1,000 × 2,000in + 400out = 2,000,000 input,400,000 output tokens

Estimated cost across 18 models

google Cheapest

Gemini 2.5 Flash-Lite

$0.360/ month
$4.32 / year
openai

GPT-6 luna

$0.400/ month
$4.80 / year
deepseek

DeepSeek V4.1-Flash

$1.08/ month
$12.96 / year
Peak hours
google

Gemini 3.1 Flash-Lite

$1.10/ month
$13.20 / year
google

Gemini 2.5 Flash

$2.70/ month
$32.40 / year
Long context ≥200,000
google

Gemini 3.8 Flash

$3.00/ month
$36.00 / year
anthropic

Claude Haiku 4.5

$4.00/ month
$48.00 / year
deepseek

DeepSeek V4-Pro

$4.22/ month
$50.69 / year
Peak hours
grok

Grok 4.7

$6.40/ month
$76.80 / year
openai

GPT-6 sol

$8.00/ month
$96.00 / year
anthropic

Sonnet 5

$8.00/ month
$96.00 / year
google

Gemini 3.1 Pro Preview

$8.80/ month
$106 / year
google

Gemini 2.5 Pro

$11.00/ month
$132 / year
Long context ≥200,000
openai

GPT-5.6 sol

$16.00/ month
$192 / year
anthropic

Opus 5.5

$16.00/ month
$192 / year
openai

GPT-6 astra

$40.00/ month
$480 / year
anthropic

Fable 5.1

$40.00/ month
$480 / year
openai

GPT-5.6 cyber

$55.00/ month
$660 / year

Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json

Quick takeaways
  • 1Output tokens usually cost 4–8× more than input — keeping replies short is the highest-leverage cost saver.
  • 2Anthropic prompt caching drops repeat-input cost by ~90%; OpenAI cached input drops it by ~90% on GPT-5.
  • 3DeepSeek off-peak hours (16:30–00:30 UTC) are 10× cheaper than peak — schedule batch jobs accordingly.
  • 4Gemini 2.5 Pro doubles its input price above 200K context — large-context calls cost roughly twice as much.

OpenRouter markup, explained

OpenRouter charges a small markup on top of each provider's official API price. The current average is around 5%, but the exact markup varies per model and per credit package. For most single-model workloads the overhead is small; it gets painful when you pay per request with many small models, or when you don't need OpenRouter's fallback / multi-model routing at all.

Current average markup: ×1.05 (5%)

When going direct is cheaper

If your workload is single-model, high-volume, and latency-sensitive, going direct to the provider's API almost always wins: lower latency, no aggregator in the request path, and you can negotiate volume pricing the aggregator can't match. OpenRouter wins when you need (a) automatic fallback across providers, (b) one bill for many models, or (c) access to models that aren't available on your home region.

Token Cost Calculator

How to bring the bill down

  1. 1Use a smaller model for routing / classification — Haiku 4.5 / GPT-5 nano cost roughly 5–10× less than the flagship.
  2. 2Cache repeat prefixes (system prompts, tool schemas) — Anthropic cached input drops the rate by ~90%.
  3. 3Set max_tokens explicitly — output is the expensive side; capping it prevents runaway replies.
  4. 4Batch non-urgent jobs into DeepSeek off-peak hours (16:30–00:30 UTC) for ~10× cheaper inference.
  5. 5Trim long context — Gemini 2.5 Pro doubles its input rate above 200K; chunk documents that don't need the rest.
Token Cost Calculator

calculator.faqTitle

OpenRouter charges the provider's official API price plus a small markup (typically 5%) for the routing, billing, and fallback service. The exact markup varies per model — check the model's card on OpenRouter for the current number.