DeepSeek V4 Flash pricing & benchmarks
Released Apr 24, 2026 · served by 21 providers · ranked #28 by intelligence in our index
What DeepSeek V4 Flash is
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
- Model ID:
deepseek/deepseek-v4-flash - Modality: text->text
- Tool calling: supported · Structured output: supported
- Weights:
deepseek-ai/DeepSeek-V4-Flashon Hugging Face
Providers and prices for DeepSeek V4 Flash
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| StreamLakecheapest | $0.090 | $0.179 | 1.0M | fp8 | 95.5% | — |
| DeepInfra | $0.090 | $0.180 | 1.0M | fp4 | 99.0% | — |
| Baidu | $0.091 | $0.182 | 1.0M | fp8 | 99.9% | — |
| GMICloud | $0.094 | $0.188 | 1.0M | fp8 | 98.7% | — |
| AkashML | $0.098 | $0.196 | 131K | fp8 | 99.3% | — |
| DigitalOcean | $0.112 | $0.224 | 262K | — | 97.5% | — |
| Alibaba | $0.134 | $0.268 | 1M | — | 99.8% | — |
| Venice | $0.138 | $0.275 | 1M | — | 98.7% | — |
| Morph | $0.139 | $0.278 | 1.0M | — | 96.6% | — |
| Io Net | $0.139 | $0.278 | 33K | fp8 | 98.9% | — |
| SiliconFlow | $0.130 | $0.280 | 1.0M | fp8 | 97.9% | — |
| Ionstream | $0.140 | $0.280 | 1.0M | fp4 | 91.8% | — |
| Parasail | $0.140 | $0.280 | 1.0M | fp8 | 89.8% | — |
| Fireworks | $0.140 | $0.280 | 1.0M | — | 95.9% | — |
| Novita | $0.140 | $0.280 | 1.0M | fp8 | 99.7% | — |
| Ambient | $0.140 | $0.280 | 1.0M | fp4 | 81.7% | — |
| Cloudflare | $0.140 | $0.280 | 384K | — | 99.6% | — |
| AtlasCloud | $0.140 | $0.280 | 1.0M | fp4 | 99.3% | — |
| DeepSeek | $0.140 | $0.280 | 1.0M | — | 99.9% | — |
| CoreWeave | $0.140 | $0.280 | 1.0M | fp8 | 97.4% | — |
| Mancer 2 | $0.250 | $1.00 | 1.0M | fp8 | 93.2% | — |
DeepSeek V4 Flash benchmark results
Design Arena head-to-head results, as reported through the OpenRouter model API.
| Category | Arena | Rank | Elo | Win rate |
|---|---|---|---|---|
| svg | models | #26 | 1204 | 48.5% |
| gamedev | models | #32 | 1253 | 50.3% |
| 3d | models | #35 | 1246 | 49.4% |
| codecategories | models | #37 | 1237 | 49.3% |
| website | models | #38 | 1233 | 49.7% |
| asciiart | models | #44 | 1153 | 43.1% |
| uicomponent | models | #51 | 1200 | 44.9% |
| dataviz | models | #67 | 1155 | 40.7% |
Cheaper models in the same class
Models scoring within 4 points of DeepSeek V4 Flash on the Intelligence Index, but with a lower output price.
FAQ
How much does the DeepSeek V4 Flash API cost?
$0.090 per 1M input tokens and $0.179 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.14.
Is DeepSeek V4 Flash free?
No. It is a paid model starting at $0.090 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for DeepSeek V4 Flash?
StreamLake at $0.090 input / $0.179 output per 1M tokens (fp8 quantization).
Can I self-host DeepSeek V4 Flash?
Yes — weights are published as deepseek-ai/DeepSeek-V4-Flash on Hugging Face, so you can run it on your own hardware.