StackProofStackProof

Model pricing · DeepSeek · value tier · August 2026

DeepSeek-V4-Flash API pricing, in numbers that matter

List price, real cost per task, rank against the field, and the self-hosting break-even, all computed from our open dataset.

Input, $/1M tokens

$0.44

Output, $/1M tokens

$1.32

Cost per task

$0.00396

Cost rank

#2 of 10

The task profile behind every figure: 1,500 input and 500 output tokens per call, 3 calls per task. On that profile DeepSeek-V4-Flash costs $0.00396 per task and $3,960 per million tasks. Output is billed at 3x the input rate, so despite the profile sending three times more input than output, output makes up 50% of the bill. Input and output carry comparable weight here, so prompt caching and context trimming return nearly as much as output caps.

DeepSeek-V4-Flash against the field

#Model$/taskvs DeepSeek-V4-FlashHead-to-head
1Mistral Small$0.001580.40xcompare
2DeepSeek-V4-Flash$0.00396--
3Gemini 2.5 Flash$0.005101.29xcompare
4Grok 4.3$0.009382.37xcompare
5GPT-5.4 mini$0.010132.56xcompare
6Claude Haiku 4.5$0.012003.03xcompare
7Claude Sonnet 5$0.024006.06xcompare
8Gemini 3.1 Pro$0.027006.82xcompare
9Claude Opus 4.8$0.0600015.2xcompare
10GPT-5.5$0.0675017.0xcompare

Self-hosting break-even

Against a reference $2/hour dedicated instance ($1,460/month fixed, 788,400 tasks/month capacity at full utilisation), the break-even against DeepSeek-V4-Flash is 368,687 tasks per month. That sits inside the box's capacity, so past that volume self-hosting genuinely wins on cost. Work through your own numbers in the $/task calculator or the break-even guide.

Frequently asked

How much does the DeepSeek-V4-Flash API cost?
DeepSeek-V4-Flash lists at $0.44 per million input tokens and $1.32 per million output tokens as of August 2026. On a realistic task of 1,500 input and 500 output tokens across 3 calls, that is $0.00396 per task, or $3,960 per million tasks.
Is DeepSeek-V4-Flash expensive compared to other models?
DeepSeek-V4-Flash ranks #2 of 10 tracked models on cost per task. It costs 2.51x more than the cheapest tracked model (Mistral Small), while the most expensive (GPT-5.5) costs 17.0x more than it.
When does self-hosting beat the DeepSeek-V4-Flash API?
Against a $2/hour dedicated GPU ($1,460/month), self-hosting becomes cheaper than DeepSeek-V4-Flash above 368,687 tasks per month, within that box's physical capacity of 788,400 tasks.

See also: all tracked models · every head-to-head comparison · download the dataset