Llama Guard 4 12B pricing & benchmarks
Released Apr 30, 2025 · served by 2 providers
What Llama Guard 4 12B is
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
- Model ID:
meta-llama/llama-guard-4-12b - Modality: text+image->text
- Knowledge cutoff: 2024-08-31
- Tool calling: not supported · Structured output: not supported
- Weights:
meta-llama/Llama-Guard-4-12Bon Hugging Face
Providers and prices for Llama Guard 4 12B
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.180 | $0.180 | 164K | bf16 | 99.4% | — |
| Together | $0.200 | $0.200 | 1.0M | — | 99.0% | — |
FAQ
How much does the Llama Guard 4 12B API cost?
$0.180 per 1M input tokens and $0.180 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.23.
Is Llama Guard 4 12B free?
No. It is a paid model starting at $0.180 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Llama Guard 4 12B?
DeepInfra at $0.180 input / $0.180 output per 1M tokens (bf16 quantization).
Can I self-host Llama Guard 4 12B?
Yes — weights are published as meta-llama/Llama-Guard-4-12B on Hugging Face, so you can run it on your own hardware.