GPT-6 and GPT-5 per-token rates, Batch API 50% off, and prompt-cache savings — compared to Claude, Gemini, and DeepSeek.
Last updated: 2026-09-26
Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json
OpenAI's Batch API processes requests asynchronously with a 24-hour SLA at half the standard rate — both input and output bill at 50%. You submit a JSONL file of requests, poll for completion, and there are no rate limits on the queue. Ideal for non-urgent workloads: evals, bulk summarization, dataset generation, and large backfills.
50% off · 24-hour SLA · no queue rate-limits
OpenAI automatically caches prompt prefixes (≥1,024 tokens) for 5–10 minutes (longer on GPT-5/6). Cached reads bill at roughly 10% of the standard input rate. For repeat-prefix workloads (system prompts, tool schemas, large documents) caching is the single biggest cost lever — most production stacks see a 50–80% input reduction once caching is on.
Compare prices across OpenAI, Claude, Gemini, DeepSeek, and Grok. Real tokenizer for OpenAI, accurate estimates for the rest.
OpenSee what your tokens actually cost after the OpenRouter markup — and when going direct to the provider is cheaper.
OpenCost per token and monthly estimates for Claude Opus, Sonnet, and Haiku — with the Anthropic API.
OpenV4-Pro and V4.1-Flash per-token rates, off-peak hours, and prompt-cache savings — compared to Claude, GPT, and Gemini.
Open