Nemotron 3 Super pricing & benchmarks
Released Mar 11, 2026 · served by 3 providers
What Nemotron 3 Super is
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
- Model ID:
nvidia/nemotron-3-super-120b-a12b - Modality: text->text
- Tool calling: supported · Structured output: supported
- Weights:
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8on Hugging Face
Providers and prices for Nemotron 3 Super
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.085 | $0.400 | 262K | bf16 | 91.0% | — |
| DigitalOcean | $0.210 | $0.455 | 1M | — | 96.4% | — |
| Nebius | $0.300 | $0.900 | 262K | fp4 | 99.7% | — |
Cheaper models in the same class
Models scoring within 4 points of Nemotron 3 Super on the Intelligence Index, but with a lower output price.
FAQ
How much does the Nemotron 3 Super API cost?
$0.085 per 1M input tokens and $0.400 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $0.20.
Is Nemotron 3 Super free?
No. It is a paid model starting at $0.085 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Nemotron 3 Super?
DeepInfra at $0.085 input / $0.400 output per 1M tokens (bf16 quantization).
Can I self-host Nemotron 3 Super?
Yes — weights are published as nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 on Hugging Face, so you can run it on your own hardware.