google Open weights Vision Reasoning

Gemma 4 26B A4B pricing & benchmarks

Released Apr 3, 2026 · served by 9 providers

Input / 1M tokens$0.070
Output / 1M tokens$0.300
Context window262K
Cached input$0.050
Intelligence index25.7
Coding index39.3
Agentic index11.0
1M in + 300K out$0.16

What Gemma 4 26B A4B is

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Providers and prices for Gemma 4 26B A4B

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Cloudflarecheapest $0.100 $0.300 256K 99.4%
DeepInfra $0.070 $0.340 262K fp8 99.7%
SiliconFlow $0.120 $0.400 262K fp8 93.3%
Novita $0.130 $0.400 262K bf16 99.6%
Ionstream $0.130 $0.400 262K bf16 92.3%
Parasail $0.130 $0.400 262K bf16 98.9%
Venice $0.130 $0.400 256K bf16 98.6%
NextBit $0.140 $0.420 262K bf16 98.1%
Google $0.150 $0.600 262K 91.9%

Cheaper models in the same class

Models scoring within 4 points of Gemma 4 26B A4B on the Intelligence Index, but with a lower output price.

FAQ

How much does the Gemma 4 26B A4B API cost?

$0.070 per 1M input tokens and $0.300 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.16.

Is Gemma 4 26B A4B free?

No. It is a paid model starting at $0.070 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for Gemma 4 26B A4B ?

Cloudflare at $0.100 input / $0.300 output per 1M tokens.

Can I self-host Gemma 4 26B A4B ?

Yes — weights are published as google/gemma-4-26B-A4B-it on Hugging Face, so you can run it on your own hardware.