StackProofStackProof

Model pricing · Google · fast tier · August 2026

Gemini 2.5 Flash API pricing, in numbers that matter

List price, real cost per task, rank against the field, and the self-hosting break-even, all computed from our open dataset.

Input, $/1M tokens

$0.3

Output, $/1M tokens

$2.5

Cost per task

$0.00510

Cost rank

#3 of 10

The task profile behind every figure: 1,500 input and 500 output tokens per call, 3 calls per task. On that profile Gemini 2.5 Flash costs $0.00510 per task and $5,100 per million tasks. Output is billed at 8.33x the input rate, so despite the profile sending three times more input than output, output makes up 74% of the bill. Capping response length is the highest-return prompt-level optimisation on this model.

Gemini 2.5 Flash against the field

#Model$/taskvs Gemini 2.5 FlashHead-to-head
1Mistral Small$0.001580.31xcompare
2DeepSeek-V4-Flash$0.003960.78xcompare
3Gemini 2.5 Flash$0.00510--
4Grok 4.3$0.009381.84xcompare
5GPT-5.4 mini$0.010131.99xcompare
6Claude Haiku 4.5$0.012002.35xcompare
7Claude Sonnet 5$0.024004.71xcompare
8Gemini 3.1 Pro$0.027005.29xcompare
9Claude Opus 4.8$0.0600011.8xcompare
10GPT-5.5$0.0675013.2xcompare

Self-hosting break-even

Against a reference $2/hour dedicated instance ($1,460/month fixed, 788,400 tasks/month capacity at full utilisation), the break-even against Gemini 2.5 Flash is 286,275 tasks per month. That sits inside the box's capacity, so past that volume self-hosting genuinely wins on cost. Work through your own numbers in the $/task calculator or the break-even guide.

Frequently asked

How much does the Gemini 2.5 Flash API cost?
Gemini 2.5 Flash lists at $0.3 per million input tokens and $2.5 per million output tokens as of August 2026. On a realistic task of 1,500 input and 500 output tokens across 3 calls, that is $0.00510 per task, or $5,100 per million tasks.
Is Gemini 2.5 Flash expensive compared to other models?
Gemini 2.5 Flash ranks #3 of 10 tracked models on cost per task. It costs 3.24x more than the cheapest tracked model (Mistral Small), while the most expensive (GPT-5.5) costs 13.2x more than it.
When does self-hosting beat the Gemini 2.5 Flash API?
Against a $2/hour dedicated GPU ($1,460/month), self-hosting becomes cheaper than Gemini 2.5 Flash above 286,275 tasks per month, within that box's physical capacity of 788,400 tasks.

See also: all tracked models · every head-to-head comparison · download the dataset