NVIDIA Unveils Nemotron 3 Ultra: America's Largest and Smartest Open-Weight Model
On June 1, 2026, during the Computex conference in Taipei, NVIDIA unveiled Nemotron 3 Ultra, its largest and most powerful open-weight artificial intelligence model ever created. With a total of 550 billion parameters and 55 billion active parameters, this model represents a significant leap forward in the global competition for the development of artificial intelligence models.
Architecture and Performance
Nemotron 3 Ultra uses a "mixture-of-experts" architecture, a design that allows only the relevant parameters for a specific request to be activated, reducing operational costs and improving efficiency. This approach enables the model to perform inferences up to five times faster than Chinese rivals and with 30% lower operational costs than comparable open-weight alternatives.
The model was evaluated by Artificial Analysis, an independent entity that awarded Nemotron 3 Ultra a score of 48 on the Intelligence Index, a composite benchmark that aggregates 10 evaluations ranging from reasoning, programming, general knowledge, and agentic performance. This score positions Nemotron 3 Ultra as the smartest open-weight model produced in the United States, far surpassing national competitors like Google's Gemma 4 31B and OpenAI's gpt-oss-120b.
The Nemotron Family
NVIDIA announced the Nemotron family in November 2023, with the launch of the third generation in December 2025. The Nemotron family includes three models of different sizes: Nano for light tasks, Super for mid-level business applications, and Ultra for complex reasoning workloads. All three models share the same hybrid architecture that combines Mamba-2 layers, standard Transformer attention, and mixture-of-experts routing.
Mamba-2 is an alternative to standard attention that processes long sequences at a fraction of the cost, relevant for models capable of keeping a million tokens in memory simultaneously. Nemotron 3 Ultra supports a context window of 1 million tokens, theoretically allowing an agent to have an entire large codebase or hundreds of research documents at its disposal simultaneously.
The Ultra model also includes a technique called multi-token prediction (MTP), which allows the model to predict several future tokens simultaneously instead of one at a time, accelerating generation. All three Nemotron 3 models have been post-trained using reinforcement learning in multiple interactive environments, teaching them to plan and execute multi-step tasks rather than just answering questions.
Accessibility and Speed
Nemotron 3 Ultra is publicly available, and its training recipes will be made public. To run it, a supercomputer is essential, but it can be accessed via NVIDIA's API or cloud service providers without owning the necessary hardware, similar to using models like GPT or Claude through a browser.
The speed of Nemotron 3 Ultra is one of its strengths. On a pre-release DeepInfra endpoint, the model served over 300 output tokens per second. Chinese models in the same intelligence class, such as DeepSeek V4 Pro and Kimi K2.6, were served at 50–100 tokens per second through their current commercial APIs. This difference in speed is crucial for practical implementations, especially for autonomous agents performing complex, multi-step tasks where waiting for each step quickly adds up.
The Global Competition
Despite its impressive performance, Nemotron 3 Ultra fails to surpass Chinese models in terms of intelligence. For example, Moonshot AI's Kimi K2.6 has a score of 54 on the Intelligence Index, six points higher than Nemotron 3 Ultra. This gap represents a significant difference, positioning Kimi K2.6 in fourth place among all global artificial intelligence models, both open and closed, just three points behind the leading proprietary models from Anthropic, Google, and OpenAI.
The current situation reflects a trend where Chinese labs are flooding the open ecosystem with strong models, while American companies like OpenAI, Anthropic, and Google keep their best systems behind APIs. NVIDIA is the largest American name actively seeking to reverse this trend, with a five-year plan to spend $26 billion on the development of open-weight artificial intelligence models.
The Future of Nemotron
Nemotron 3 Ultra is the most visible result of this investment so far. NVIDIA has also announced that it is already working on Nemotron 4, the next generation, developed through the Nemotron Coalition, a group of eight artificial intelligence labs, including Mistral AI and Perplexity, assembled by NVIDIA in March 2026 to co-develop open frontier models on DGX Cloud infrastructure. Nemotron 3 Ultra will be available from June 4, 2026.
The Impact of Nemotron 3 Ultra on the Global Market
The announcement of Nemotron 3 Ultra represents a crucial moment for the artificial intelligence industry, marking a significant step forward in the development of open-weight models. This model's advanced capabilities and accessibility through NVIDIA's API or cloud service providers have the potential to reshape the market landscape.
Implications for the Market and Users
The availability of Nemotron 3 Ultra through NVIDIA's API or cloud service providers makes this model accessible to a wide range of users, even without the need for specialized hardware. This could have a significant impact on the market, allowing developers and companies to integrate advanced artificial intelligence models into their applications without incurring the high costs of a supercomputer.
However, the speed and efficiency of Nemotron 3 Ultra are not sufficient to surpass Chinese models in terms of intelligence. This means that while NVIDIA continues to improve its technologies, users seeking the highest performance may still need to consider models like Kimi K2.6 for applications requiring high reasoning and problem-solving capabilities.
Nemotron 3 Ultra represents an important step for NVIDIA and the United States in the global competition for open-weight artificial intelligence. However, the technological gap with Chinese models remains significant. The road to the future of open-weight AI will be long and complex, but with strategic investments and international collaborations, NVIDIA may be able to reverse the trend and bring the United States back to a leadership position in the sector.
Frequently Asked Questions
- What is the main difference between Nemotron 3 Ultra and Chinese models? The main difference lies in the Intelligence Index, where Moonshot AI's Kimi K2.6 surpasses Nemotron 3 Ultra by six points, positioning it in fourth place in the global ranking.
- How can I access Nemotron 3 Ultra? Nemotron 3 Ultra is available through NVIDIA's API or cloud service providers, allowing users to access it without specialized hardware.
- What are NVIDIA's future plans for open-weight AI? NVIDIA has announced a five-year, $26 billion plan for the development of open-weight artificial intelligence models and is already working on Nemotron 4, the next generation of models.
Editorial Note and Disclaimer
The guides and content published on GoYou are the result of independent research and analysis activities, for informational, educational, and in-depth purposes.
GoYou does not constitute a journalistic publication or an editorial product pursuant to Law No. 62/2001 and does not engage in real-time information activities.
The GoYou project does not provide professional, technical, legal, or financial advice and disclaims any liability for the improper use of the information published.
In the Crypto sector, every investment involves risks: readers are invited to always inform themselves autonomously before making any decision.