V4-Pro and V4.1-Flash per-token rates, off-peak hours, and prompt-cache savings — compared to Claude, GPT, and Gemini.
Last updated: 2026-09-26
Prices verified 2026-09-26. Source: official provider pricing pages. · Machine-readable: /api/prices.json
Mon–Fri (excluding Chinese public holidays), DeepSeek bills at the full peak rate only during 01:00–04:00 and 06:00–10:00 UTC. All other times — including all weekend hours — are off-peak and bill at 50% of peak for both input and output.
Off-peak = 50% of peak · Mon–Fri · weekends & Chinese holidays all-day off-peak
DeepSeek cache hits are aggressive — both V4.1-Flash and V4-Pro drop cached input by ~99% from the standard rate. For any workload with a repeated prefix (system prompt, tool definitions, document context) caching the prefix is the single highest-leverage cost saving.
Compare prices across OpenAI, Claude, Gemini, DeepSeek, and Grok. Real tokenizer for OpenAI, accurate estimates for the rest.
OpenSee what your tokens actually cost after the OpenRouter markup — and when going direct to the provider is cheaper.
OpenCost per token and monthly estimates for Claude Opus, Sonnet, and Haiku — with the Anthropic API.
OpenGPT-6 and GPT-5 per-token rates, Batch API 50% off, and prompt-cache savings — compared to Claude, Gemini, and DeepSeek.
Open