Gemma 4 31B pricing & benchmarks
Released Apr 2, 2026 · served by 16 providers · ranked #57 by intelligence in our index
What Gemma 4 31B is
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
- Model ID:
google/gemma-4-31b-it - Modality: text+image+video->text
- Tool calling: supported · Structured output: supported
- Weights:
google/gemma-4-31B-iton Hugging Face
Providers and prices for Gemma 4 31B
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.090 | $0.340 | 262K | fp4 | 96.6% | — |
| CoreWeave | $0.100 | $0.340 | 262K | bf16 | 95.0% | — |
| OpenInference | $0.100 | $0.350 | 262K | bf16 | 98.3% | — |
| Venice | $0.120 | $0.360 | 256K | bf16 | 99.7% | — |
| Chutes | $0.120 | $0.370 | 131K | fp4 | 89.9% | — |
| DeepInfra | $0.130 | $0.380 | 262K | fp8 | 93.7% | — |
| SiliconFlow | $0.130 | $0.400 | 262K | fp8 | 83.3% | — |
| Novita | $0.140 | $0.400 | 262K | bf16 | 62.5% | — |
| Friendli | $0.140 | $0.400 | 262K | — | 99.1% | — |
| Morph | $0.140 | $0.400 | 175K | fp4 | 98.0% | — |
| Crusoe | $0.140 | $0.400 | 262K | — | 98.3% | — |
| Parasail | $0.150 | $0.400 | 262K | fp8 | 97.4% | — |
| Phala | $0.150 | $0.460 | 262K | — | 96.9% | — |
| ModelRun | $0.220 | $0.550 | 262K | fp4 | 95.2% | — |
| Together | $0.390 | $0.970 | 262K | — | 96.3% | — |
| SambaNova | $0.380 | $1.15 | 131K | — | 98.8% | — |
| Cerebras | $0.990 | $1.49 | 131K | fp16 | 100.0% | — |
Cheaper models in the same class
Models scoring within 4 points of Gemma 4 31B on the Intelligence Index, but with a lower output price.
FAQ
How much does the Gemma 4 31B API cost?
$0.090 per 1M input tokens and $0.340 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.19.
Is Gemma 4 31B free?
No. It is a paid model starting at $0.090 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Gemma 4 31B?
DeepInfra at $0.090 input / $0.340 output per 1M tokens (fp4 quantization).
Can I self-host Gemma 4 31B?
Yes — weights are published as google/gemma-4-31B-it on Hugging Face, so you can run it on your own hardware.