meta-llama Open weights

Llama 3.1 8B Instruct pricing & benchmarks

Released Jul 23, 2024 · served by 5 providers

Input / 1M tokens$0.020
Output / 1M tokens$0.040
Context window131K
Cached input$0.025
Intelligence index7.6
Coding index5.4
Agentic index0.5
1M in + 300K out$0.03

What Llama 3.1 8B Instruct is

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Providers and prices for Llama 3.1 8B Instruct

The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.

ProviderInputOutputContext QuantUptime 24hThroughput
DeepInfracheapest $0.020 $0.040 131K fp8 99.7%
Novita $0.020 $0.050 16K fp8 99.3%
Groq $0.050 $0.080 131K 100.0%
CoreWeave $0.220 $0.220 128K bf16 99.8%
Cloudflare $0.152 $0.287 32K fp8 99.6%

Cheaper models in the same class

Models scoring within 4 points of Llama 3.1 8B Instruct on the Intelligence Index, but with a lower output price.

FAQ

How much does the Llama 3.1 8B Instruct API cost?

$0.020 per 1M input tokens and $0.040 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.03.

Is Llama 3.1 8B Instruct free?

No. It is a paid model starting at $0.020 per 1M input tokens, though some providers offer trial credits.

Which provider is cheapest for Llama 3.1 8B Instruct?

DeepInfra at $0.020 input / $0.040 output per 1M tokens (fp8 quantization).

Can I self-host Llama 3.1 8B Instruct?

Yes — weights are published as meta-llama/Meta-Llama-3.1-8B-Instruct on Hugging Face, so you can run it on your own hardware.