Nemotron 3 Ultra pricing & benchmarks
Released Jun 4, 2026 · served by 4 providers · ranked #36 by intelligence in our index
What Nemotron 3 Ultra is
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
- Model ID:
nvidia/nemotron-3-ultra-550b-a55b - Modality: text->text
- Tool calling: supported · Structured output: supported
- Weights:
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16on Hugging Face
Providers and prices for Nemotron 3 Ultra
The same model costs different amounts depending on who serves it. Prices are per 1M tokens, cheapest first. Uptime and throughput are OpenRouter's rolling measurements, not vendor claims.
| Provider | Input | Output | Context | Quant | Uptime 24h | Throughput |
|---|---|---|---|---|---|---|
| DeepInfracheapest | $0.500 | $2.20 | 262K | fp4 | 94.9% | — |
| BaseTen | $0.600 | $2.40 | 203K | fp4 | 98.7% | — |
| Venice | $0.625 | $3.13 | 256K | fp8 | 91.0% | — |
| Together | $0.600 | $3.60 | 512K | — | 96.6% | — |
Nemotron 3 Ultra benchmark results
Design Arena head-to-head results, as reported through the OpenRouter model API.
| Category | Arena | Rank | Elo | Win rate |
|---|---|---|---|---|
| svg | models | #49 | 1128 | 37.6% |
| 3d | models | #50 | 1190 | 41.8% |
| asciiart | models | #50 | 1112 | 37.5% |
| uicomponent | models | #59 | 1180 | 39.5% |
| gamedev | models | #60 | 1183 | 38.8% |
| codecategories | models | #69 | 1155 | 36.2% |
| dataviz | models | #71 | 1151 | 37.4% |
| website | models | #82 | 1127 | 32.5% |
Cheaper models in the same class
Models scoring within 4 points of Nemotron 3 Ultra on the Intelligence Index, but with a lower output price.
FAQ
How much does the Nemotron 3 Ultra API cost?
$0.500 per 1M input tokens and $2.20 per 1M output tokens. A typical workload of 1M input + 300K output tokens costs about $1.16.
Is Nemotron 3 Ultra free?
No. It is a paid model starting at $0.500 per 1M input tokens, though some providers offer trial credits.
Which provider is cheapest for Nemotron 3 Ultra?
DeepInfra at $0.500 input / $2.20 output per 1M tokens (fp4 quantization).
Can I self-host Nemotron 3 Ultra?
Yes — weights are published as nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 on Hugging Face, so you can run it on your own hardware.