nvidia
NVIDIA Nemotron 3 Super 120B A12B
- Input / 1M
- $0.1575
- Output / 1M
- $0.6825
- Context
- 256,000 tokens
- Catalog
- Gateway catalog
About this model
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. It delivers up to 7x higher throughput, providing fast, cost-efficient inference for agentic tasks. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere.
Use an endpoint supported by this model.
Get a key for this setup →const response = await fetch('https://api.minirouter.sh/v1/chat/completions', { method: 'POST', headers: { Authorization: `Bearer ${process.env.MINIROUTER_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'nvidia/nemotron-3-super-120b-a12b', messages: [{ role: 'user', content: 'Why is the sky blue?' }], max_tokens: 1024, }),}) const data = await response.json()console.log(data.choices[0].message.content)Overview
- Model type
- language
- Context window
- 256,000
- Maximum output
- 32,000
- Input / 1M tokens
- $0.1575
- Output / 1M tokens
- $0.6825
- Released
- 2026-03-11
Prices include the MiniRouter fee. Compatibility snapshot: 2026-08-06. Input: text. Output: text.
Endpoints
/v1/chat/completions/v1/messages/v1/responsesPlatform-funded routing only. Caller BYOK credentials and OIDC-based upstream authentication are intentionally excluded; authenticate with a MiniRouter bearer key.
Captured parameters: max_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputs.
Structured output: JSON Schema and JSON mode enforced on verified providers. How it works
API
Use the setup panel above to switch between Chat Completions, Messages, and client configurations. Every snippet keeps the model ID unchanged.
https://api.minirouter.sh/v1nvidia/nemotron-3-super-120b-a12bProviders
MiniRouter selects from the configured routes below. Priority is the catalog order, not a latency or availability score.
| Provider | Priority | Input / 1M | Output / 1M | Parameters | Access |
|---|---|---|---|---|---|
vercel | 1 | $0.1575 | $0.6825 | max_tokens, temperature, stop, tools, tool_choice, reasoning, include_reasoning, response_format, structured_outputs | MiniRouter balance |