Llama 4 Maverick pricing & benchmarks
Released Apr 5, 2025 · served by 5 providers
What Llama 4 Maverick is
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
- Model ID:
meta-llama/llama-4-maverick - Modality: text+image->text
- Knowledge cutoff: 2024-08-31
- Tool calling: supported · Structured output: supported
- Weights:
meta-llama/Llama-4-Maverick-17B-128E-Instructon Hugging Face
Providers and prices for Llama 4 Maverick
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.200 | $0.800 | 1.0M | fp8 | 99.8% | — |
| Novita | $0.270 | $0.850 | 1.0M | fp8 | 99.8% | — |
| DigitalOcean | $0.250 | $0.870 | 128K | — | 99.8% | — |
| Parasail | $0.350 | $1.00 | 524K | fp8 | 99.9% | — |
| $0.350 | $1.15 | 524K | — | 99.9% | — |
Llama 4 Maverick benchmark results
Design Arena head-to-head results, as reported through the OpenRouter model API.
| Category | Arena | Rank | Elo | Win rate |
|---|---|---|---|---|
| 3d | models | #99 | 957 | 40.2% |
| uicomponent | models | #103 | 936 | 40.8% |
| dataviz | models | #108 | 912 | 38.4% |
| gamedev | models | #109 | 894 | 33.7% |
| codecategories | models | #111 | 910 | 35.8% |
| website | models | #114 | 896 | 34.4% |
Cheaper models in the same class
Models scoring within 4 points of Llama 4 Maverick on the Intelligence Index, but with a lower output price.
FAQ
How much does the Llama 4 Maverick API cost?
$0.200 per 1M input tokens and $0.800 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.44.
Is Llama 4 Maverick free?
No. It is a paid model starting at $0.200 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Llama 4 Maverick?
DeepInfra at $0.200 input / $0.800 output per 1M tokens (fp8 quantization).
Can I self-host Llama 4 Maverick?
Yes — weights are published as meta-llama/Llama-4-Maverick-17B-128E-Instruct on Hugging Face, so you can run it on your own hardware.