Gemma 3 4B pricing & benchmarks
Released Mar 13, 2025 · served by 1 provider
What Gemma 3 4B is
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
- Model ID:
google/gemma-3-4b-it - Modality: text+image->text
- Knowledge cutoff: 2024-08-31
- Tool calling: not supported · Structured output: supported
- Weights:
google/gemma-3-4b-iton Hugging Face
Providers and prices for Gemma 3 4B
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.050 | $0.100 | 131K | bf16 | 100.0% | — |
FAQ
How much does the Gemma 3 4B API cost?
$0.050 per 1M input tokens and $0.100 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.08.
Is Gemma 3 4B free?
No. It is a paid model starting at $0.050 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Gemma 3 4B?
DeepInfra at $0.050 input / $0.100 output per 1M tokens (bf16 quantization).
Can I self-host Gemma 3 4B?
Yes — weights are published as google/gemma-3-4b-it on Hugging Face, so you can run it on your own hardware.