Token Cost Calculator

ChatGPT API Pricing

GPT-6 and GPT-5 per-token rates, Batch API 50% off, and prompt-cache savings — compared to Claude, Gemini, and DeepSeek.

Last updated: 2026-09-26

1,000
2,000
400
1,000 × 2,000in + 400out = 2,000,000 input,400,000 output tokens

Estimated cost across 18 models

google Cheapest

Gemini 2.5 Flash-Lite

$0.360/ month
$4.32 / year
openai

GPT-6 luna

$0.400/ month
$4.80 / year
deepseek

DeepSeek V4.1-Flash

$1.08/ month
$12.96 / year
Peak hours
google

Gemini 3.1 Flash-Lite

$1.10/ month
$13.20 / year
google

Gemini 2.5 Flash

$2.70/ month
$32.40 / year
Long context ≥200,000
google

Gemini 3.8 Flash

$3.00/ month
$36.00 / year
anthropic

Claude Haiku 4.5

$4.00/ month
$48.00 / year
deepseek

DeepSeek V4-Pro

$4.22/ month
$50.69 / year
Peak hours
grok

Grok 4.7

$6.40/ month
$76.80 / year
openai

GPT-6 sol

$8.00/ month
$96.00 / year
anthropic

Sonnet 5

$8.00/ month
$96.00 / year
google

Gemini 3.1 Pro Preview

$8.80/ month
$106 / year
google

Gemini 2.5 Pro

$11.00/ month
$132 / year
Long context ≥200,000
openai

GPT-5.6 sol

$16.00/ month
$192 / year
anthropic

Opus 5.5

$16.00/ month
$192 / year
openai

GPT-6 astra

$40.00/ month
$480 / year
anthropic

Fable 5.1

$40.00/ month
$480 / year
openai

GPT-5.6 cyber

$55.00/ month
$660 / year

Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json

Quick takeaways
  • 1GPT-5 cached input drops to $0.125/M — a 90% discount vs the $1.25/M standard rate.
  • 2The Batch API runs at 50% off with a 24-hour SLA — ideal for evals, bulk summarization, and dataset generation.
  • 3GPT-6 astra is the most expensive active model at $10 input / $50 output per 1M tokens.
  • 4GPT-6 luna is the cheapest active model at $0.10 input / $0.50 output per 1M tokens.

Batch API · 50% off

OpenAI's Batch API processes requests asynchronously with a 24-hour SLA at half the standard rate — both input and output bill at 50%. You submit a JSONL file of requests, poll for completion, and there are no rate limits on the queue. Ideal for non-urgent workloads: evals, bulk summarization, dataset generation, and large backfills.

50% off · 24-hour SLA · no queue rate-limits

Prompt caching · ~90% off

OpenAI automatically caches prompt prefixes (≥1,024 tokens) for 5–10 minutes (longer on GPT-5/6). Cached reads bill at roughly 10% of the standard input rate. For repeat-prefix workloads (system prompts, tool schemas, large documents) caching is the single biggest cost lever — most production stacks see a 50–80% input reduction once caching is on.

  • GPT-6 astra: $1.000/M (90% off)
  • GPT-6 sol: $0.200/M (90% off)
  • GPT-6 luna: $0.010/M (90% off)
Token Cost Calculator

How to bring the bill down

  1. 1Use GPT-5 nano or GPT-6 luna for routing / classification — output is roughly 8–20× cheaper than the flagship.
  2. 2Enable cached input — OpenAI auto-caches prefix repeats for ~90% off the standard input rate.
  3. 3Route non-urgent jobs through the Batch API for 50% off (24-hour SLA, no queue rate-limits).
  4. 4Set max_tokens explicitly — output is the expensive side; capping it prevents runaway replies.
  5. 5ChatGPT Plus ($20/mo) does NOT replace API access — high-volume workloads still need the API.
Token Cost Calculator

calculator.faqTitle

GPT-6 astra is $10 input / $50 output per 1M; GPT-6 sol is $2 / $10; GPT-6 luna is $0.10 / $0.50. Older GPT-5 (now delisted) was $1.25 / $10 with cached input at $0.125/M. Use the calculator above to compare every active model at your exact usage.