Cheapest capable LLMs for high-volume work (2026)
What is the lowest cost per million tokens without dropping to a weak model?
Rebuilt from live model data · August 3, 2026
How we ranked
Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers.
Top 3 for high-volume work
- 1
DeepSeek V4 Flash
Output / 1M: $0.179 $0.090 in / $0.179 out per 1M 1.0M context 21 providersDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Open weights Tool calling Reasoning Released Apr 24, 2026 - 2
Hy3 preview
Output / 1M: $0.210 $0.063 in / $0.210 out per 1M 262K context 1 providerHy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...
Open weights Tool calling Reasoning Released Apr 22, 2026 - 3
MiMo-V2.5
Output / 1M: $0.224 $0.112 in / $0.224 out per 1M 1.1M context 7 providersMiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Open weights Vision Tool calling Reasoning Released Apr 22, 2026
Full ranking
| # | Model | Output / 1M | Output / 1M | Context | Providers | Weights |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash | $0.179 | $0.179 | 1.0M | 21 | Open |
| 2 | Hy3 preview | $0.210 | $0.210 | 262K | 1 | Open |
| 3 | MiMo-V2.5 | $0.224 | $0.224 | 1.1M | 7 | Open |
| 4 | DeepSeek V3.2 | $0.311 | $0.311 | 164K | 14 | Open |
| 5 | GPT-5.4 Nano | $0.625 | $0.625 | 400K | 2 | Closed |
| 6 | Ring-2.6-1T | $0.625 | $0.625 | 262K | 1 | Closed |
| 7 | MiMo-V2.5-Pro | $0.696 | $0.696 | 1.1M | 6 | Open |
| 8 | DeepSeek V4 Pro | $0.870 | $0.870 | 1.0M | 18 | Open |
| 9 | DeepSeek V3.1 Terminus | $0.950 | $0.950 | 164K | 5 | Open |
| 10 | MiniMax M3 | $0.960 | $0.960 | 1.0M | 8 | Open |
| 11 | MiniMax M2.7 | $0.960 | $0.960 | 205K | 10 | Open |
| 12 | Nex-N2-Pro | $1.00 | $1.00 | 262K | 2 | Open |
FAQ
What is the lowest cost per million tokens without dropping to a weak model?
DeepSeek V4 Flash leads for high-volume work with output / 1m $0.179, at $0.179 per 1M output tokens. Hy3 preview ($0.210) and MiMo-V2.5 ($0.224) follow. Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers.
How is this ranking produced?
Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers. The list rebuilds from the live model catalogue on every deploy, so a model that launches or changes price appears here without an editor rewriting the page. Last rebuild: August 3, 2026.
What does the top pick cost to run?
DeepSeek V4 Flash costs about $0.14 for a workload of 1M input + 300K output tokens, based on the cheapest of 21 provider(s).