z-ai Open weights Vision Reasoning

GLM 4.6V pricing & benchmarks

Released Dec 8, 2025 · served by 2 providers

Input / 1M tokens$0.300
Output / 1M tokens$0.900
Context window131K
Cached input$0.055
1M in + 300K out$0.57

What GLM 4.6V is

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

Providers and prices for GLM 4.6V

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Novitacheapest $0.300 $0.900 131K bf16 99.5%
Z.AI $0.300 $0.900 131K fp8 97.7%

FAQ

How much does the GLM 4.6V API cost?

$0.300 per 1M input tokens and $0.900 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.57.

Is GLM 4.6V free?

No. It is a paid model starting at $0.300 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for GLM 4.6V?

Novita at $0.300 input / $0.900 output per 1M tokens (bf16 quantization).

Can I self-host GLM 4.6V?

Yes — weights are published as zai-org/GLM-4.6V on Hugging Face, so you can run it on your own hardware.