deepseek Open weights Reasoning

R1 Distill Llama 70B pricing & benchmarks

Released Jan 23, 2025 · served by 1 provider

Input / 1M tokens$0.800
Output / 1M tokens$0.800
Context window8K
Cached input
1M in + 300K out$1.04

What R1 Distill Llama 70B is

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

Providers and prices for R1 Distill Llama 70B

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
Novitacheapest $0.800 $0.800 8K bf16 100.0%

FAQ

How much does the R1 Distill Llama 70B API cost?

$0.800 per 1M input tokens and $0.800 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.04.

Is R1 Distill Llama 70B free?

No. It is a paid model starting at $0.800 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for R1 Distill Llama 70B?

Novita at $0.800 input / $0.800 output per 1M tokens (bf16 quantization).

Can I self-host R1 Distill Llama 70B?

Yes — weights are published as deepseek-ai/DeepSeek-R1-Distill-Llama-70B on Hugging Face, so you can run it on your own hardware.