Nvidia Nemotron 3 Ultra Tops US Open Models but Trails China

Wednesday, June 3, 2026

Nvidia unveiled Nemotron 3 Ultra at its GTC Taipei keynote on Monday (June 1, 2026) and released a benchmark, run with the evaluation firm Artificial Analysis, that places the model first among open-weight systems built in the United States and behind the Chinese-led frontier. The 550-billion-parameter model scores 48 on the Artificial Analysis Intelligence Index, ahead of Google’s Gemma 4 31B at 39 and OpenAI’s gpt-oss-120b at 33, and six points behind Moonshot’s Kimi K2.6 at 54.

Nemotron 3 Ultra is the centerpiece of a free enterprise Agent Toolkit, and Nvidia is competing to make Nvidia hardware the default place enterprise AI agents run. The model is free; the speed and runtime bundled around it are tuned to Nvidia silicon.

Model Specifications

The model uses a mixture-of-experts design, with roughly 550 billion total parameters but about 55 billion active per token. On a pre-release endpoint at DeepInfra it served more than 300 tokens per second, against the 50 to 100 that comparably sized models from DeepSeek and Moonshot manage in the market today.

Nvidia’s headline claims Nemotron runs up to five times faster and up to 30% cheaper than open frontier rivals in its class. Nvidia has disclosed a five-year, $26 billion plan to fund open-weight development, and says Nemotron 4 is already in progress.

Agent Toolkit Components

Nvidia paired the model with three other parts of the Agent Toolkit released Monday:

  • NemoClaw — open framework for building the orchestration layer that turns a model into an agent. Available now.
  • OpenShell — secure runtime that sets privacy and policy controls. Early preview.
  • CUDA-X skills — Nvidia libraries (cuDF, cuOpt, etc.) exposed to agents as reusable skills under open licenses such as Apache 2.0, but CUDA- and GPU-dependent.

“NVIDIA NemoClaw provides enterprise software developers with the open building blocks to create more secure, long-running AI coworkers that amplify human expertise as they reshape how work gets done,” Jensen Huang said.

Named Customers and Partners

  • Nvidia itself is the first customer using Cadence’s ChipStack agent to verify its own chip designs.
  • CrowdStrike runs agents on Nemotron models to identify and remediate software vulnerabilities.
  • Palantir folded the models into air-gapped Forward Deployed Engineer systems.
  • Siemens and Synopsys use NemoClaw for chip-design workflows.
  • Foxconn is building manufacturing agent MoMClaw on the stack.

Availability

Nemotron 3 Ultra is expected to reach Hugging Face, OpenRouter and build.nvidia.com on June 4, 2026 as an Nvidia NIM microservice. NemoClaw is available now and OpenShell is in early preview.

Goldman Sachs analyst James Schneider kept a buy rating with a $285 price target, noting Nvidia is “aggressively investing to drive the adoption of agentic AI across developers and ecosystem partners.”