NVIDIA Nemotron 3 Ultra
NVIDIA's largest Nemotron 3 reasoning model for complex agents, coding, planning, tools, RAG, and million-token analysis.
NVIDIA's largest Nemotron 3 reasoning model for complex agents, coding, planning, tools, RAG, and million-token analysis.
Plain-English overview
Nemotron 3 Ultra is a 550-billion-parameter mixture-of-experts model that activates 55 billion parameters for each token. It is built for demanding agent and reasoning work and supports up to one million tokens of text context with a configurable reasoning budget.
NVIDIA offers a prototype API, NIM deployment, partner endpoints, and downloadable weights. That range lets teams trade convenience for infrastructure control, but it also means performance and cost are properties of the chosen deployment rather than the checkpoint alone.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Example: 6,000 input tokens for 10 pages, plus a 500-token summary. Page lengths vary; adjust the numbers below.
One run sends your input to the model once and receives an answer. Tokens are pieces of text: input is what you send, output is the answer you receive.
NVIDIA Nemotron 3 Ultra
NVIDIA
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through a free NVIDIA API prototype, partner endpoints, NVIDIA NIM deployment, and downloadable weights.
NVIDIA lists global deployment for the model; actual hosted regions and self-host location depend on the endpoint or infrastructure selected.
Check live availabilityData and training
Downloaded weights can run in infrastructure the operator controls. NVIDIA's hosted prototype and partner endpoints are governed by their own trial, logging, retention, and deployment terms.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
NVIDIA's largest Nemotron 3 reasoning model for complex agents, coding, planning, tools, RAG, and million-token analysis. Nemotron 3 Ultra is a 550-billion-parameter mixture-of-experts model that activates 55 billion parameters for each token. It is built for demanding agent and reasoning work and supports up to one million tokens of text context with a configurable reasoning budget.
NVIDIA Nemotron 3 Ultra was released on June 4, 2026 according to the cited provider materials.
Available through a free NVIDIA API prototype, partner endpoints, NVIDIA NIM deployment, and downloadable weights. The access routes listed in this guide are NVIDIA NIM and Hugging Face.
Free NVIDIA prototype endpoint; production and self-host cost varies. NVIDIA offers a prototype NIM endpoint plus partner and downloadable deployment routes. Production hosting and GPU costs depend on the chosen route.
Available through a free NVIDIA API prototype, partner endpoints, NVIDIA NIM deployment, and downloadable weights. NVIDIA lists global deployment for the model; actual hosted regions and self-host location depend on the endpoint or infrastructure selected.
Downloaded weights can run in infrastructure the operator controls. NVIDIA's hosted prototype and partner endpoints are governed by their own trial, logging, retention, and deployment terms. The policy belongs to the provider route and account terms, so verify it again before production use.
One price per model, using the same input/output mix. We weight each API rate by its share of one million total tokens and use the lowest matching reviewed route. This is a comparison rate, not a cost per task or a quality score. Tokenizers, caching, reasoning and long context can change your actual bill.
75% input + 25% output · no cache discounts
OpenAI
Mistral AI
OpenAI
Anthropic
OpenAI
Anthropic
OpenAI
OpenAI
Anthropic
OpenAI
Anthropic
Anthropic
OpenAI
Anthropic
Related comparisons
A Claude model for large code changes, detailed research and business tasks that take many steps.
Open full comparisonA lower-cost GPT-5.6 option for coding, analysis and assistants that use tools.
Open full comparisonThe budget GPT-5.6 tier for processing many short, repeatable tasks.
Open full comparisonThe original Fable model for complex coding and multi-step work, retained for existing integrations.
Open full comparisonA Claude model for substantial code changes, reasoning and business workflows with many steps.
Open full comparisonA lower-priced Claude option for everyday coding, writing and tool-assisted work.
Open full comparisonA compact Claude model for quick replies, classification and narrowly scoped assistants.
Open full comparisonOpenAI's most capable model for difficult end-to-end work, combining frontier reasoning with a million-token context window and a broad set of agent tools.
Open full comparisonOpenAI's flagship general model for difficult coding, analysis, and professional work, with a very large context window and a broad native tool set.
Open full comparisonAnthropic's frontier model for ambitious, long-running agentic and coding work, with strong vision and enterprise marketplace availability.
Open full comparisonGoogle's stable, high-efficiency multimodal model for agents, software work, and large mixed-media inputs at an introductory Flash-tier price.
Open full comparisonSpaceXAI's frontier text-and-image model for coding, agentic tasks, and knowledge work, with direct web, X, and code tools.
Open full comparisonA specialist Grok model that sends several AI agents to investigate a difficult question in parallel, then combines their work into one researched answer.
Open full comparisonThinking Machines Lab's large open-weights model for customizable reasoning, coding, tools, vision, and audio workflows.
Open full comparisonMistral's open-weight frontier model for demanding multimodal, coding, reasoning, and agent workflows, with a 256K context window.
Open full comparisonMistral's lower-cost open model that combines normal instruction following, reasoning, coding, vision, and agent tools in one endpoint.
Open full comparisonAlibaba's frontier Qwen model for long-context reasoning, coding, tools, and understanding text, images, and video.
Open full comparisonMiniMax's million-token multimodal model for coding, agents, computer use, and long projects at a low direct API price.
Open full comparisonZ.ai's open-weight long-context reasoning model for coding, tools, and agentic work across self-hosted and managed routes.
Open full comparisonMoonshot AI's frontier multimodal model for million-token coding, reasoning, knowledge work, and tool-driven agents.
Open full comparisonTII's compact open reasoning model for text tasks, long contexts, and self-hosted function-calling workflows.
Open full comparisonTII's 34B open instruction model for multilingual text, coding, and controlled self-hosted applications.
Open full comparisonMeta's latest multimodal reasoning model for long-running agents, coding, tools, and complex user collaboration.
Open full comparisonCeleris' agent-focused diffusion language model for fast reasoning, tool loops, structured actions, and OpenAI-compatible integration.
Open full comparisonNVIDIA's compact 30B mixture-of-experts model for efficient specialist agents and high-volume text workflows.
Open full comparisonAlibaba's efficient million-context multimodal model for fast agents, coding, document work, vision, video understanding, and tool use.
Open full comparisonDeepSeek's flagship million-context text model for difficult reasoning, coding, long-running agents, tools, and very large outputs.
Open full comparisonDeepSeek's lower-cost V4 model for high-volume reasoning, coding, agents, and million-token text workloads.
Open full comparison