NVIDIA Nemotron 3 Super 120B A12B
NVIDIA · Active · available in 13 of 31 regions
- Context window
- —
- Max output
- —
- Input / 1M
- $0.150
- us-east-1
- Output / 1M
- $0.650
What NVIDIA Nemotron 3 Super 120B A12B is good for
editorialNVIDIA Nemotron 3 Super 120B A12B, at the premium end of this provider’s range at $0.150 per 1M input tokens. Callable in 13 of 31 regions.
Suits
- Bulk jobs. Batch inference is available, typically at half the on-demand rate, for work that can wait — backfills, nightly enrichment, evaluation runs.
Think twice if
- Agentic loops. No prompt caching published, so every step re-pays full price for the prompt prefix. Expensive for agents that resend a long context.
- Function calling. No tool use, so it cannot drive an agent loop or call your APIs — it answers from what is in the prompt.
- Documents with layout. Text only. Scanned PDFs, screenshots and charts need a vision model or an OCR step first.
- High volume. At $0.150 per 1M input tokens it sits at the expensive end of the NVIDIA range. Worth it for hard tasks, wasteful for bulk extraction.
This section is judgement, not data from AWS. It is composed from the capabilities, context window, price position and region coverage shown elsewhere on this page — so it stays in step with the daily snapshot rather than going stale.
Where teams typically use it
Chat and assistants
Conversational work where responsiveness matters as much as depth.
- · A customer-facing assistant where the first token needs to appear in under a second.
- · An in-product copilot that explains what the user is looking at and answers follow-ups.
- · A triage bot that qualifies an incoming request before routing it to the right team.
Model & inference profile IDs
Base model ID — direct on-demand invoke
Pricing by region
USD per 1M tokens| Region | Input | Output | Cache read | Cache write | Batch in | Batch out | Source |
|---|---|---|---|---|---|---|---|
| us-east-1 US East (N. Virginia) | $0.150 | $0.650 | — | — | $0.075 | $0.325 | API |
| us-east-2 US East (Ohio) | $0.150 | $0.650 | — | — | $0.075 | $0.325 | API |
| us-west-2 US West (Oregon) | $0.150 | $0.650 | — | — | $0.075 | $0.325 | API |
| eu-west-1 EU (Ireland) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| eu-west-2 EU (London) | $0.230 | $1.01 | — | — | $0.115 | $0.505 | API |
| eu-central-1 EU (Frankfurt) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| eu-north-1 EU (Stockholm) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| eu-south-1 EU (Milan) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| ap-northeast-1 Asia Pacific (Tokyo) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| ap-south-1 Asia Pacific (Mumbai) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| ap-southeast-2 Asia Pacific (Sydney) | $0.150 | $0.670 | — | — | $0.075 | $0.335 | API |
| ap-southeast-3 Asia Pacific (Jakarta) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
| sa-east-1 South America (Sao Paulo) | $0.180 | $0.780 | — | — | $0.090 | $0.390 | API |
Region availability
TPM / RPM are default account quotas from AWS Service Quotas where a model-specific limit is published. They are per-account defaults and adjustable on request.
Where this model actually runs
Under a strict residency constraint, NVIDIA Nemotron 3 Super 120B A12B can be served without the request leaving 12 of the 25 jurisdictions with a Bedrock region.
NVIDIA Nemotron 3 Super 120B A12B appears in
Common questions
- How much does NVIDIA Nemotron 3 Super 120B A12B cost on Amazon Bedrock?
- $0.150 per 1M input tokens and $0.650 per 1M output tokens in us-east-1.
- Which AWS regions support NVIDIA Nemotron 3 Super 120B A12B?
- 13 of 31 regions: us-east-1, us-east-2, us-west-2, eu-west-1, eu-west-2, eu-central-1, eu-north-1, eu-south-1, ap-northeast-1, ap-south-1, ap-southeast-2, ap-southeast-3, sa-east-1.
- What is the model ID for NVIDIA Nemotron 3 Super 120B A12B on Bedrock?
- nvidia.nemotron-super-3-120b.