z-ai Open weights Reasoning

GLM 4.7 Flash pricing & benchmarks

Released Jan 19, 2026 · served by 4 providers

Input / 1M tokens$0.060
Output / 1M tokens$0.400
Context window203K
Cached input$0.010
1M in + 300K out$0.18

What GLM 4.7 Flash is

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

Providers and prices for GLM 4.7 Flash

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
DeepInfracheapest $0.060 $0.400 203K bf16 99.6%
Cloudflare $0.060 $0.400 131K 99.6%
Novita $0.070 $0.400 200K bf16 86.4%
Venice $0.125 $0.500 128K fp8 98.6%

GLM 4.7 Flash benchmark results

Design Arena head-to-head results, as reported through the OpenRouter model API.

CategoryArenaRankEloWin rate
uicomponent models #35 1244 57.6%
website models #42 1219 54%
codecategories models #44 1208 53.1%
3d models #54 1178 51.2%
gamedev models #54 1191 49.7%
svg models #55 1088 44.2%
dataviz models #70 1152 45.3%

FAQ

How much does the GLM 4.7 Flash API cost?

$0.060 per 1M input tokens and $0.400 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.18.

Is GLM 4.7 Flash free?

No. It is a paid model starting at $0.060 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for GLM 4.7 Flash?

DeepInfra at $0.060 input / $0.400 output per 1M tokens (bf16 quantization).

Can I self-host GLM 4.7 Flash?

Yes — weights are published as zai-org/GLM-4.7-Flash on Hugging Face, so you can run it on your own hardware.