DeepSeek · DeepSeek V4
DeepSeek V4 (V4-Flash and V4-Pro)
DeepSeek's V4 API generation, offered as a cheap high-volume Flash model and a costlier, more capable Pro model, both with a million-token context window.
Treat the price as an invitation to test, not a reason to route everything here. Score it on accepted work against your current model, and confirm the data terms before any regulated or sensitive task touches it.
Use it for
- High-volume or cost-sensitive tasks where token price dominates
- Long-context document and code work up to roughly a million tokens
- Comparing an aggressively cheap option against a frontier baseline on the same task
Do not assume a low token price means a low total cost, and do not route sensitive or regulated material to it before reading DeepSeek's own data-handling terms, which are not established in the cited pricing and release documentation.
Effort without the jargon
The task is high-volume or cost-sensitive and can tolerate the cheaper model's ceiling on quality.
The task needs more careful reasoning and the roughly threefold price step over Flash is worth it.
Advanced recordCost, limits, privacy and operations
V4-Flash is cheap enough to route a lot of volume through, but DeepSeek has not published a portable data-retention rule in the cited sources, so read its terms before sending anything sensitive.
- API cost
- Officially disclosedDeepSeek prices V4-Flash at US$0.44 per million input tokens on a cache miss (US$0.014 on a cache hit) and US$1.32 per million output tokens at peak rates. Off-peak rates are half of peak, and peak is 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday. V4-Pro is three times Flash on both input and output. Work out which window your job runs in before estimating cost.
- Context window
- Officially disclosed1,000,000 tokens
- Maximum output
- Officially disclosed384,000 tokens
- Knowledge cutoff
- Not disclosedNot disclosed in the cited record
- Recorded inputs
- Text
- Recorded output
- Text · Structured JSON · Tool calls
Data handling is a product decision
Not disclosedThe cited release note and pricing page do not state a portable retention or training rule for V4. Confirm DeepSeek's API and product terms before sending sensitive or regulated material.
Operational constraints
- DeepSeek documents both thinking and non-thinking modes as configurable per request; confirm which one is the default before comparing cost.
- Billing is peak and off-peak: off-peak is half the peak rate outside 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, so the same job costs twice as much inside those windows.
- Data retention and training terms are not established in the cited pricing and release documentation, so sensitive material needs a separate terms check.
- A third model, deepseek-v4-flash-vision-exp, is listed as experimental and vision capable at the same price as V4-Flash. Experimental is DeepSeek's own label, so treat it as subject to change rather than as a settled option.
Unknown means the cited official record does not disclose a safe value. It is not an estimate. Prices are provider list prices in USD where stated and can change before this record's review date.
Keep the evidence separate
DeepSeek describes its V4 generation as two API models, V4-Flash and V4-Pro, each with a one-million-token context window and up to 384K output tokens, supporting thinking and non-thinking modes, JSON output and tool calls. V4-Flash is priced at US$0.44 per million input tokens and US$1.32 per million output at peak rates, with off-peak billing at half those rates.
No sufficiently useful independent note has been added yet. That absence is not filled with our guess.
No Human Bit result is claimed; use this as evaluation guidance only.
Not scheduled
No Human Bit result is claimed for this model yet.