Nemotron 3 Ultra Arrives on JumpStart
Visual status: no verified article image is available. The reporting remains text-first.
Nemotron 3 Ultra just hit JumpStart, delivering 5x faster inference for long running agent workflows.
AWS announced day zero availability of NVIDIA Nemotron 3 Ultra on Amazon SageMaker JumpStart. The model is described as an open large language model built for frontier reasoning and orchestration in long running autonomous agents, designed to keep planning, tool calls, and self correction loops moving across hundreds of turns without blowing through compute budgets. The Nemotron 3 Ultra package carries 550 billion total parameters, with 55 billion active parameters per forward pass, and it runs on a hybrid Transformer Mamba Mixture of Experts architecture. The model is optimized for the NVFP4 format, which the team says helps to push throughput higher and cost lower for agentic workloads. Context length goes up to 1 million tokens, and inference speed is described as 5x faster for long running agent workflows, with benchmarks indicating up to 30 percent lower cost for complex agent tasks. The one click deployment within JumpStart is the fastest route to start experimenting with frontier levels of agentic AI, according to the AWS post.
The heart of Nemotron 3 Ultra is the MoE approach. The team reports that only 55 billion of the 550 billion total parameters are active for a given forward pass, which helps sustain throughput as context grows. In practice this means agents can plan, call tools, delegate work to sub agents, check results, and keep going across hundreds of turns while the compute bill stays more manageable. The architecture is explicitly designed to support frontier intelligence at a fraction of the compute cost of dense models of equivalent quality, a critical factor for teams embedding multi-step reasoning and external tool use into production workflows. By combining a Transformer backbone with the MoE gating, Nemotron 3 Ultra aims to keep latency predictable even as tasks require long horizons and multi-turn deliberation.
From the perspective of practitioners, there are clear levers and caveats. The one click JumpStart deployment lowers the barrier to trying a model of this scale in enterprise contexts, letting product teams move from concept to pilot in weeks rather than months. The NVFP4 precision and the 1 million token context length open new possibilities for sustained agentic reasoning, where a single session might span long decision sequences and external calls. At the same time, operators should watch how the MoE routing behaves under real workloads. While the model promises lower cost per task, the actual savings depend on the mix of tasks, the distribution of tokens across calls, and the hardware employed for deployment. The Nemotron 3 Ultra design explicitly targets agentic workloads that require planning and multi-turn interaction, so teams should prepare for integration work around tool catalogs, error handling, and monitoring of long-horizon reasoning accuracy.
In practice, JumpStart day zero availability signals a shift in how enterprises prototype and iterate with frontier AI. Instead of building bespoke inference pipelines for huge models, teams can spin up experiments with a pay-as-you-go, one-click deployment and then scale once value is demonstrated. That combination of a large but selectively active parameter budget, a purpose built MoE architecture, and enterprise friendly deployment tools positions Nemotron 3 Ultra as a notable option for those chasing long context, multi-turn, tool enabled AI agents. The model's emphasis on agentic workflows, planning, tool use, and cross turn reasoning, highlights what operators should prioritize when evaluating next generation AI systems: real time throughput under heavy context, cost per task, and the reliability of multi-turn orchestration.
Sources
- NVIDIA Nemotron 3 Ultra now available on Amazon SageMaker JumpStartAWS Machine Learning / Primary source / Published JUN 04, 2026 / Accessed JUN 05, 2026
- Fundamental’s Large Tabular Model NEXUS is now available on Amazon SageMaker JumpStartAWS Machine Learning / Primary source / Published JUN 03, 2026 / Accessed JUN 05, 2026