z-ai Open weights Reasoning

GLM 4.5 Air pricing & benchmarks

Released Jul 25, 2025 · served by 3 providers

Input / 1M tokens$0.130
Output / 1M tokens$0.850
Context window131K
Cached input$0.025
1M in + 300K out$0.39

What GLM 4.5 Air is

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Providers and prices for GLM 4.5 Air

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Novitacheapest $0.130 $0.850 131K bf16 97.8%
SiliconFlow $0.140 $0.860 131K fp8 99.3%
Z.AI $0.200 $1.10 131K fp8 97.7%

GLM 4.5 Air benchmark results

Design Arena head-to-head results, as reported through the OpenRouter model API.

CategoryArenaRankEloWin rate
dataviz models #42 1222 59.4%
svg models #51 1121 50.8%
3d models #53 1180 54.1%
uicomponent models #63 1161 54.6%
website models #65 1171 51.3%
codecategories models #66 1168 51.5%
gamedev models #68 1150 48.4%

FAQ

How much does the GLM 4.5 Air API cost?

$0.130 per 1M input tokens and $0.850 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.39.

Is GLM 4.5 Air free?

No. It is a paid model starting at $0.130 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for GLM 4.5 Air?

Novita at $0.130 input / $0.850 output per 1M tokens (bf16 quantization).

Can I self-host GLM 4.5 Air?

Yes — weights are published as zai-org/GLM-4.5-Air on Hugging Face, so you can run it on your own hardware.