R1 Distill Llama 70B pricing & benchmarks
Released Jan 23, 2025 · served by 1 provider
What R1 Distill Llama 70B is
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...
- Model ID:
deepseek/deepseek-r1-distill-llama-70b - Modality: text->text
- Knowledge cutoff: 2024-07-31
- Tool calling: not supported · Structured output: not supported
- Weights:
deepseek-ai/DeepSeek-R1-Distill-Llama-70Bon Hugging Face
Providers and prices for R1 Distill Llama 70B
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| Novitacheapest | $0.800 | $0.800 | 8K | bf16 | 100.0% | — |
FAQ
How much does the R1 Distill Llama 70B API cost?
$0.800 per 1M input tokens and $0.800 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.04.
Is R1 Distill Llama 70B free?
No. It is a paid model starting at $0.800 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for R1 Distill Llama 70B?
Novita at $0.800 input / $0.800 output per 1M tokens (bf16 quantization).
Can I self-host R1 Distill Llama 70B?
Yes — weights are published as deepseek-ai/DeepSeek-R1-Distill-Llama-70B on Hugging Face, so you can run it on your own hardware.