nvidia vera rubin nvl72 agentic inference is the rack-scale system NVIDIA built to handle the messy, multi-step reality of AI agents instead of simple chatbot replies. It packs 72 Rubin GPUs and 36 Vera CPUs into one liquid-cooled domain connected by sixth-generation NVLink, delivering up to 3,600 PFLOPS of NVFP4 inference performance and 20.7 TB of pooled HBM4 memory. Early measurements on real agentic coding workloads show it can produce 30x higher throughput per megawatt than the previous-generation GB300 NVL72 while cutting cost per million tokens dramatically. For anyone running or planning agent fleets—coding assistants, research agents, multi-tool workflows—this is the platform that finally makes continuous, high-concurrency agent operation practical inside a fixed power budget.
- Combines 72 Rubin GPUs + 36 Vera CPUs in a single NVLink domain for rack-scale coherence.
- Targets agentic workloads that grow context, call tools, and spawn sub-agents over many turns.
- Delivers up to 30x more agentic work per megawatt versus GB300 NVL72 on SemiAnalysis AgentX traces.
- Cuts token costs enough that always-on agents become economically viable rather than premium experiments.
- Ships with the broader Vera Rubin platform (including optional Groq 3 LPX racks) for both high-throughput and low-latency decode.
What nvidia vera rubin nvl72 agentic inference Actually Is
Think of a traditional inference server as a fast short-order cook. An agentic workload is more like a head chef running a multi-course dinner service with constant interruptions—tool calls, context growth, sub-agent handoffs, and long reasoning chains. The cook burns out. The chef needs a kitchen designed for the whole service.
That’s the shift NVIDIA engineered into the Vera Rubin NVL72. The Rubin GPU itself carries 336 billion transistors, up to 288 GB of HBM4 per package, and a third-generation Transformer Engine tuned for NVFP4. Pair two of those GPUs with one Vera CPU (88 custom Olympus cores) over coherent NVLink-C2C and you get a superchip. Scale that to 72 GPUs and 36 CPUs inside one rack, glue them with NVLink 6 (roughly 260 TB/s of all-to-all bandwidth), and the rack behaves like one giant accelerator.
nvidia vera rubin nvl72 agentic inference The software layer matters just as much. NVIDIA Dynamo orchestrates disaggregated serving—prefill on one set of resources, decode on another—so the expensive high-bandwidth memory stays busy on the right phase of the request. Large-scale expert parallelism spreads mixture-of-experts layers across the entire NVLink domain. The result is sustained tokens-per-second on trillion-parameter models with long contexts without the usual efficiency cliff.
Why Agentic Inference Breaks Older Racks
A simple chat turn might burn a few thousand tokens. An agent session that plans, retrieves, calls tools, and reasons across multiple steps can easily consume 10–15x more, according to usage patterns reported by OpenRouter and measured in production traces. Context keeps growing. KV cache balloons. Prefill and decode fight for the same silicon.
Older platforms force a painful trade-off: either keep latency low and throughput collapses, or chase throughput and interactive agents start feeling sluggish. Power walls appear fast. Token costs climb. Operators end up rationing agent use or accepting poor economics.
nvidia vera rubin nvl72 agentic inference attacks that trade-off at the system level. Early SemiAnalysis AgentX results (real recorded coding sessions with preserved tool calls and sub-agent behavior) show the platform delivering up to 30x higher throughput per megawatt than GB300 NVL72 on models such as DeepSeek V4 Pro while hitting interactive tokens-per-second targets. MLPerf Inference v6.1 preview submissions further showed 2.5x–3.7x higher throughput on demanding models like DeepSeek-R1 and Qwen3-VL compared with the same prior rack. NVIDIA’s own numbers put the cost-per-million-tokens advantage in the range of 10x versus Blackwell-generation systems for long-context reasoning workloads, with some AgentX comparisons reaching even higher reductions.
These are not lab cherry-picks. They come from workloads that look like the ones enterprises actually run.
Key Specs Side by Side
| Metric | Vera Rubin NVL72 | GB300 NVL72 (prior gen) | Practical Difference for Agents |
|---|---|---|---|
| GPUs / CPUs | 72 Rubin + 36 Vera | 72 Blackwell + Grace | More coherent CPU orchestration for tool use |
| NVFP4 Inference | 3,600 PFLOPS | Lower (generation gap) | Higher sustained generation rate |
| HBM Capacity | 20.7 TB HBM4 | Smaller HBM3e pool | Longer contexts without aggressive offload |
| NVLink Domain | NVLink 6, ~260 TB/s | NVLink 5 | Better expert parallelism & KV distribution |
| Reported AgentX Gain | Up to 30x throughput/MW | Baseline | Far more concurrent agent sessions per rack |
| Token Cost Trend | Up to 10x–45x lower (workload dependent) | Higher | Continuous agents become affordable |
Sources: NVIDIA product specifications and published AgentX / MLPerf results as of late 2026.

Step-by-Step: How a Beginner Should Approach nvidia vera rubin nvl72 agentic inference
- Map the actual agent traffic, not the marketing FLOPS.
Record a few dozen real sessions. Note context growth, tool-call frequency, and acceptable tokens-per-second per user. That profile tells you whether you need pure throughput or a mix with low-latency decode (where Groq 3 LPX racks pair well). - Start with a single rack or cloud instance rather than a full factory.
Most major clouds and NVIDIA Cloud Partners began offering Vera Rubin capacity in the second half of 2026. Spin up a test workload with Dynamo or TensorRT-LLM / vLLM / SGLang. Measure tokens per megawatt and cost per million tokens on your own traces. - Enable disaggregated serving early.
Keep prefill and decode phases on the resources that suit them. Rate-matching between the two phases is where a lot of the efficiency appears. NVIDIA’s documentation walks through the configuration. - Right-size the KV cache and expert parallelism.
With 20.7 TB of HBM4 in the domain you have more headroom, but still plan the cache hierarchy. Offload cold context to the BlueField-4 storage path when sessions idle. - Instrument power and utilization at the rack level.
Intelligent Power Smoothing and liquid cooling let operators push higher sustained utilization inside the same facility power envelope. Track that number religiously—it is the real capacity metric now. - Pilot one production agent fleet, then expand.
Once the economics clear (lower cost per successful agent task), scale the number of concurrent sessions. That is the point where the 30x throughput-per-megawatt claim starts paying the power bill.
What I’d do if I were evaluating this tomorrow: run the exact same AgentX-style workload (or my own production traces) on both a current Blackwell rack and a Vera Rubin instance side by side. Numbers, not slides, decide the purchase.
Common Mistakes & How to Fix Them
Treating it like a bigger GPU.
The win comes from the full stack—CPU orchestration, NVLink domain, disaggregated serving, and networking. Ignoring Dynamo or expert parallelism leaves performance on the table. Fix: follow the reference serving recipes NVIDIA publishes for agentic models.
Ignoring tool-calling latency.
Early AgentX numbers did not fully include Vera CPU acceleration of tool loops. If your agents make heavy external calls, measure end-to-end latency, not just pure generation. Fix: keep the Vera CPUs in the critical path for orchestration.
Over-provisioning for peak instead of sustained.
Agent traffic is bursty but long-lived. Design for average concurrent sessions at the target tokens-per-second rather than the absolute peak of a single request. Fix: use the rack-level power and thermal telemetry to set realistic concurrency limits.
Skipping software maturity checks.
Preview MLPerf and AgentX results improve with later software stacks. Confirm the framework versions and quantization paths (NVFP4, MXFP4/8) you plan to run in production. Fix: lock a software baseline and re-benchmark after each major framework update.
Where the Platform Fits in a Real AI Factory
nvidia vera rubin nvl72 agentic inference sits at the center of a larger POD design. Pair it with Spectrum-X Ethernet or Quantum-X800 InfiniBand for scale-out, BlueField-4 for storage and security offload, and optional Groq 3 LPX racks when you need deterministic sub-100 ms decode on the largest models. The whole system is liquid-cooled, cable-free modular, and designed so operators can service trays without tearing the rack apart.
For U.S. data-center operators the power story is the practical one. Facilities already constrained on megawatts can extract far more useful agent work from the same electrical and cooling budget. That is the difference between running a few pilot agents and running thousands of always-on agents across development, support, research, and operations teams.
Key Takeaways
- nvidia vera rubin nvl72 agentic inference is a 72-GPU + 36-CPU rack purpose-built for multi-step agent workloads, not single-turn chat.
- Real-world AgentX measurements show up to 30x higher throughput per megawatt versus the prior NVL72 generation.
- Token costs drop enough that continuous agent operation moves from experiment to production economics.
- Disaggregated serving, expert parallelism, and the coherent NVLink domain are the techniques that unlock the gains.
- Pair with Groq 3 LPX when interactive latency on trillion-parameter models is non-negotiable.
- Start with workload traces, not FLOPS charts; measure tokens per megawatt on your own traffic.
- Software (Dynamo, TensorRT-LLM, vLLM) and power telemetry matter as much as the silicon.
- Cloud and on-prem options both exist; evaluate them the same way—on sustained agent throughput inside your power envelope.
The real benefit is simple: more agent work for the same power and lower cost per successful task. That is what lets teams move from occasional agent demos to always-on agent fleets. Next step—pull a representative set of your current agent sessions, request a Vera Rubin trial instance from a cloud partner or NVIDIA, and run the comparison yourself. The numbers will tell you whether the upgrade pays for itself inside one budget cycle.
FAQs
How does nvidia vera rubin nvl72 agentic inference differ from a standard GPU server for agents?
It treats the entire rack as one coherent domain with high-bandwidth all-to-all communication, dedicated CPU orchestration, and software that separates prefill from decode. Ordinary servers force every phase onto the same limited memory and interconnect, which collapses under growing context and tool use.
Is nvidia vera rubin nvl72 agentic inference available in the United States right now?
Yes. Full production began in 2026. Major U.S. cloud providers and NVIDIA Cloud Partners started offering capacity in the second half of the year, with on-prem racks shipping through the MGX ecosystem.
What should I measure first when testing nvidia vera rubin nvl72 agentic inference?
Tokens per second per concurrent user at your target latency, total tokens delivered per megawatt, and fully loaded cost per million tokens on real agent traces that include tool calls and context growth. Those three numbers decide whether the platform improves your unit economics.