Cheapest capable LLMs for high-volume work (2026)

What is the lowest cost per million tokens without dropping to a weak model?

Rebuilt from live model data · August 3, 2026

How we ranked

Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers.

Top 3 for high-volume work

  1. 1

    DeepSeek V4 Flash

    Output / 1M: $0.179 $0.090 in / $0.179 out per 1M 1.0M context 21 providers

    DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

    Open weights Tool calling Reasoning Released Apr 24, 2026
  2. 2

    Hy3 preview

    Output / 1M: $0.210 $0.063 in / $0.210 out per 1M 262K context 1 provider

    Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

    Open weights Tool calling Reasoning Released Apr 22, 2026
  3. 3

    MiMo-V2.5

    Output / 1M: $0.224 $0.112 in / $0.224 out per 1M 1.1M context 7 providers

    MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

    Open weights Vision Tool calling Reasoning Released Apr 22, 2026

Full ranking

#ModelOutput / 1MOutput / 1M ContextProvidersWeights
1 DeepSeek V4 Flash $0.179 $0.179 1.0M 21 Open
2 Hy3 preview $0.210 $0.210 262K 1 Open
3 MiMo-V2.5 $0.224 $0.224 1.1M 7 Open
4 DeepSeek V3.2 $0.311 $0.311 164K 14 Open
5 GPT-5.4 Nano $0.625 $0.625 400K 2 Closed
6 Ring-2.6-1T $0.625 $0.625 262K 1 Closed
7 MiMo-V2.5-Pro $0.696 $0.696 1.1M 6 Open
8 DeepSeek V4 Pro $0.870 $0.870 1.0M 18 Open
9 DeepSeek V3.1 Terminus $0.950 $0.950 164K 5 Open
10 MiniMax M3 $0.960 $0.960 1.0M 8 Open
11 MiniMax M2.7 $0.960 $0.960 205K 10 Open
12 Nex-N2-Pro $1.00 $1.00 262K 2 Open

FAQ

What is the lowest cost per million tokens without dropping to a weak model?

DeepSeek V4 Flash leads for high-volume work with output / 1m $0.179, at $0.179 per 1M output tokens. Hy3 preview ($0.210) and MiMo-V2.5 ($0.224) follow. Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers.

How is this ranking produced?

Models scoring at least 30 on the Intelligence Index, ranked by cheapest output price per 1M tokens across all providers. The list rebuilds from the live model catalogue on every deploy, so a model that launches or changes price appears here without an editor rewriting the page. Last rebuild: August 3, 2026.

What does the top pick cost to run?

DeepSeek V4 Flash costs about $0.14 for a workload of 1M input + 300K output tokens, based on the cheapest of 21 provider(s).