google Open weights Vision

Gemma 3 4B pricing & benchmarks

Released Mar 13, 2025 · served by 1 provider

Input / 1M tokens$0.050
Output / 1M tokens$0.100
Context window131K
Cached input
Coding index2.7
1M in + 300K out$0.08

What Gemma 3 4B is

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Providers and prices for Gemma 3 4B

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
DeepInfracheapest $0.050 $0.100 131K bf16 100.0%

FAQ

How much does the Gemma 3 4B API cost?

$0.050 per 1M input tokens and $0.100 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.08.

Is Gemma 3 4B free?

No. It is a paid model starting at $0.050 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for Gemma 3 4B?

DeepInfra at $0.050 input / $0.100 output per 1M tokens (bf16 quantization).

Can I self-host Gemma 3 4B?

Yes — weights are published as google/gemma-3-4b-it on Hugging Face, so you can run it on your own hardware.