Nemotron 3 Ultra now available on AI Gateway
Nvidia's Nemotron 3 Ultra now ships on Vercel AI Gateway — 1M context, 350 tokens/sec, 30% cheaper for agents.

Why it matters
Nvidia's open reasoning model is now accessible to developers via Vercel's unified API layer, lowering friction for teams building multi-turn agent workflows. This represents a shift toward standardized model access infrastructure commoditizing agent deployment.
The key facts
9 to knowNemotron 3 Ultra: 1M token context window
Mixture-of-Experts architecture for reasoning & agent orchestration
350 tokens per second throughput
30% lower cost on agentic tasks vs. alternatives
Targets multi-turn workflows: planning, tool use, sub-agent delegation, error recovery
Available via Vercel AI Gateway with unified API, usage tracking, failover, dynamic provider sorting
AI Gateway reflects provider pricing with zero platform fee on inference
BYOK (Bring Your Own Key) support included
Published June 4, 2026
Go to the source
Vercel Blogvercel.com
Publisher excerpt: Nemotron 3 Ultra from Nvidia is now available on .Vercel AI Gateway Nemotron 3 Ultra is an open Mixture-of-Experts reasoning model built for orchestrating long-running agent workflows, with a 1M token context window. The model targets multi-turn agent workflows: planning, tool use, sub-agent…