z-ai Open weights Vision Reasoning

GLM 4.5V pricing & benchmarks

Released Aug 11, 2025 · served by 2 providers

Input / 1M tokens$0.600
Output / 1M tokens$1.80
Context window66K
Cached input$0.110
1M in + 300K out$1.14

What GLM 4.5V is

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,...

Providers and prices for GLM 4.5V

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Novitacheapest $0.600 $1.80 66K fp8 98.8%
Z.AI $0.600 $1.80 66K fp8 98.5%

FAQ

How much does the GLM 4.5V API cost?

$0.600 per 1M input tokens and $1.80 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.14.

Is GLM 4.5V free?

No. It is a paid model starting at $0.600 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for GLM 4.5V?

Novita at $0.600 input / $1.80 output per 1M tokens (fp8 quantization).

Can I self-host GLM 4.5V?

Yes — weights are published as zai-org/GLM-4.5V on Hugging Face, so you can run it on your own hardware.