Skip to main content
The Human Bit
← All signals
DDeepSeek✓ ReviewedTest carefullyChecked 3 September 2026

DeepSeek V4 is aggressively cheap at a million tokens. Cheap is not the only column.

DeepSeek's V4 generation offers a one-million-token context and 384K-token output at a fraction of frontier pricing (V4-Flash at $0.44 in and $1.32 out per million at peak, half that off peak), with thinking modes and tool calls.

Bit, the Human Bit guide

Bit’s takeaway

What not to assumeDo not assume a low token price means a low total cost, and do not route sensitive material anywhere before reading the data-handling terms.

What changed

DeepSeek released its V4 generation with two API models. V4-Flash and V4-Pro both carry a one-million-token context window and up to 384K output tokens, support thinking and non-thinking modes, JSON output and tool calls; Flash is priced at US$0.44 per million input tokens (US$0.014 on cache hits) and US$1.32 per million output at peak rates, with off-peak billing at half those rates; confirm Pro pricing against DeepSeek's current page before costing a pipeline.

Why it matters

Pricing this far below the frontier makes 'run it on everything' financially thinkable, which is exactly when the non-price columns matter: where your data goes, what the provider retains, and whether the quality holds on your task rather than on a benchmark.

Who should care

  • Technical teams with high-volume, low-risk AI workloads
  • People comparing long-context options on cost per accepted result

What to do

Trial it only on workloads whose data you would be comfortable seeing handled under the provider's terms, which you should read first. Score cost per accepted result against your current route, not cost per token, and keep anything sensitive on infrastructure you have vetted.

The human take

Tools change fast. Your judgment matters more.

A person reads the terms, decides which data may travel, and judges the model on accepted work rather than the price list.

Affected guidance

Put this to work

Verified facts

  • DeepSeek released its V4 API generation: V4-Flash and V4-Pro with one-million-token context and up to 384K output tokens, thinking and non-thinking modes and tool calls, with Flash priced at US$0.44 per million input tokens and US$1.32 per million output at peak, half those rates off peak, and US$0.014 per million on a cache hit.Checked

Sources and method

Checked
Review due
Owner
The Human Bit editorial

The Human Bit Weekly

The useful changes, not every launch.

One short issue every Monday. What changed, what it means for your work, and the part that stays yours.

The Human Bit records when and how consent was given. Subscription is confirmed only after the email provider accepts the request.