Hermes 4 405B pricing & benchmarks
Released Aug 26, 2025 · served by 1 provider
What Hermes 4 405B is
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
- Model ID:
nousresearch/hermes-4-405b - Modality: text->text
- Knowledge cutoff: 2024-08-31
- Tool calling: not supported · Structured output: not supported
- Weights:
NousResearch/Hermes-4-405Bon Hugging Face
Providers and prices for Hermes 4 405B
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| Nebiuscheapest | $1.00 | $3.00 | 131K | fp8 | 100.0% | — |
FAQ
How much does the Hermes 4 405B API cost?
$1.00 per 1M input tokens and $3.00 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.90.
Is Hermes 4 405B free?
No. It is a paid model starting at $1.00 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Hermes 4 405B?
Nebius at $1.00 input / $3.00 output per 1M tokens (fp8 quantization).
Can I self-host Hermes 4 405B?
Yes — weights are published as NousResearch/Hermes-4-405B on Hugging Face, so you can run it on your own hardware.