Token Cost Calculator

DeepSeek API Pricing

V4-Pro and V4.1-Flash per-token rates, off-peak hours, and prompt-cache savings — compared to Claude, GPT, and Gemini.

Last updated: 2026-09-26

1,000
2,000
400
1,000 × 2,000in + 400out = 2,000,000 input,400,000 output tokens

Estimated cost across 18 models

google Cheapest

Gemini 2.5 Flash-Lite

$0.360/ month
$4.32 / year
openai

GPT-6 luna

$0.400/ month
$4.80 / year
deepseek

DeepSeek V4.1-Flash

$1.08/ month
$12.96 / year
Peak hours
google

Gemini 3.1 Flash-Lite

$1.10/ month
$13.20 / year
google

Gemini 2.5 Flash

$2.70/ month
$32.40 / year
Long context ≥200,000
google

Gemini 3.8 Flash

$3.00/ month
$36.00 / year
anthropic

Claude Haiku 4.5

$4.00/ month
$48.00 / year
deepseek

DeepSeek V4-Pro

$4.22/ month
$50.69 / year
Peak hours
grok

Grok 4.7

$6.40/ month
$76.80 / year
openai

GPT-6 sol

$8.00/ month
$96.00 / year
anthropic

Sonnet 5

$8.00/ month
$96.00 / year
google

Gemini 3.1 Pro Preview

$8.80/ month
$106 / year
google

Gemini 2.5 Pro

$11.00/ month
$132 / year
Long context ≥200,000
openai

GPT-5.6 sol

$16.00/ month
$192 / year
anthropic

Opus 5.5

$16.00/ month
$192 / year
openai

GPT-6 astra

$40.00/ month
$480 / year
anthropic

Fable 5.1

$40.00/ month
$480 / year
openai

GPT-5.6 cyber

$55.00/ month
$660 / year

Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json

Quick takeaways
  • 1DeepSeek V4.1-Flash is $0.30 input / $1.20 output per 1M tokens — roughly 8× cheaper than GPT-5 on output.
  • 2V4-Pro is $1.32 / $3.96 per 1M — comparable to Claude Sonnet 4.5 on benchmarks, ~2.3× cheaper on output.
  • 3Off-peak (Mon–Fri outside 01:00–04:00 & 06:00–10:00 UTC, plus all weekends & Chinese holidays) is 50% of peak — schedule batch jobs accordingly.
  • 4Cache hits drop V4.1-Flash input from $0.30/M to $0.003/M — 99% off. Cache aggressively when the system prompt or tool schema repeats.

Off-peak hours, explained

Mon–Fri (excluding Chinese public holidays), DeepSeek bills at the full peak rate only during 01:00–04:00 and 06:00–10:00 UTC. All other times — including all weekend hours — are off-peak and bill at 50% of peak for both input and output.

Off-peak = 50% of peak · Mon–Fri · weekends & Chinese holidays all-day off-peak

Prompt-cache pricing

DeepSeek cache hits are aggressive — both V4.1-Flash and V4-Pro drop cached input by ~99% from the standard rate. For any workload with a repeated prefix (system prompt, tool definitions, document context) caching the prefix is the single highest-leverage cost saving.

  • DeepSeek V4.1-Flash: $0.003/M (99% off)
  • DeepSeek V4-Pro: $0.022/M (98% off)
Token Cost Calculator

How to bring the bill down

  1. 1Use V4.1-Flash for routing / classification — output is $1.20/M vs $3.96/M on V4-Pro, ~3.3× cheaper.
  2. 2Cache repeat prefixes — V4.1-Flash cache hit is $0.003/M, a 99% drop from $0.30/M. The single biggest cost saver for repeat-prompt workloads.
  3. 3Batch non-urgent jobs into off-peak hours (Mon–Fri outside 01:00–04:00 and 06:00–10:00 UTC) for ~50% cheaper inference.
  4. 4Set max_tokens explicitly — output is the expensive side; capping it prevents runaway replies.
  5. 5Going direct to DeepSeek's API beats OpenRouter by a few percent. OpenRouter wins on automatic fallback + a single invoice.
Token Cost Calculator

calculator.faqTitle

DeepSeek V4-Pro is $1.32 per 1M input tokens / $3.96 per 1M output tokens. DeepSeek V4.1-Flash is $0.30 / $1.20. New accounts get a small free credit to start; top-up minimums start at $5. Numbers above are re-verified against platform.deepseek.com/api-docs/pricing.