z-ai Open weights Reasoning

GLM 4.5 pricing & benchmarks

Released Jul 25, 2025 · served by 1 provider

Input / 1M tokens$0.600
Output / 1M tokens$2.20
Context window131K
Cached input$0.110
1M in + 300K out$1.26

What GLM 4.5 is

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

Providers and prices for GLM 4.5

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Z.AIcheapest $0.600 $2.20 131K fp8 100.0%

GLM 4.5 benchmark results

Design Arena head-to-head results, as reported through the OpenRouter model API.

CategoryArenaRankEloWin rate
3d models #38 1227 59.7%
svg models #44 1146 50.8%
gamedev models #49 1201 54.4%
codecategories models #55 1196 54.4%
dataviz models #55 1189 53.3%
website models #57 1195 53.8%
uicomponent models #58 1181 55.1%

FAQ

How much does the GLM 4.5 API cost?

$0.600 per 1M input tokens and $2.20 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.26.

Is GLM 4.5 free?

No. It is a paid model starting at $0.600 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for GLM 4.5?

Z.AI at $0.600 input / $2.20 output per 1M tokens (fp8 quantization).

Can I self-host GLM 4.5?

Yes — weights are published as zai-org/GLM-4.5 on Hugging Face, so you can run it on your own hardware.