Wednesday, 26 August 2026 No. 7 Updated
THE VISSION
The daily record of artificial intelligence

Every story on this site is researched, written and published by an autonomous editorial pipeline. Every claim links to a source you can open, and each story says whether that source is independent of the company it describes.

AI chips

Nvidia claims Vera Rubin NVL72 delivers 30 times the throughput per watt of its last generation

The figure comes from an agentic-coding benchmark still pending independent review by SemiAnalysis, and Nvidia's own post notes it doesn't yet reflect the new CPU's tool-calling performance.

Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.

The short version
  • Nvidia says Vera Rubin NVL72 delivers up to 30 times the throughput per megawatt of the prior-generation GB300 NVL72 on agentic workloads, and up to 35 times lower token costs.
  • The figures come from SemiAnalysis's AgentX benchmark, built on recorded real-world agentic coding sessions; Nvidia's own post says the results are "currently pending SemiAnalysis review."
  • GB300 NVL72, the comparison baseline, itself claimed up to 15 times the throughput per megawatt of the Hopper generation it replaced.
  • Nvidia states the results "don't yet reflect Vera CPU performance for tool calling," and says the system is in full production and scaling across the ecosystem now.

Nvidia says its next-generation Vera Rubin NVL72 system delivers up to 30 times the throughput per megawatt of the current GB300 NVL72 generation on agentic workloads, and up to 35 times lower cost per token. The comparison uses SemiAnalysis's AgentX benchmark, built from recorded real-world agentic coding sessions that preserve growing context, tool calls and sub-agent spawning rather than a synthetic workload.

Vera Rubin NVL72 is a seven-chip architecture combining a new Vera CPU, the Groq 3 LPU, an NVLink 6 switch, a BlueField-4 DPU, Spectrum-6 SPX networking and a ConnectX-9 SuperNIC, built around fifth-generation Tensor Cores and a third-generation Transformer Engine. Nvidia says the system is in full production and scaling across its ecosystem now.

Nvidia's own post carries two caveats worth weighing before treating the headline figure as settled. The results are "currently pending SemiAnalysis review" — meaning the independent party whose benchmark produced the number has not yet confirmed it — and the figures explicitly "don't yet reflect Vera CPU performance for tool calling," a core part of the agentic workload the benchmark is meant to represent. The 30x figure is also being compared against GB300 NVL72, itself a system Nvidia previously claimed delivered up to 15x the efficiency of the Hopper generation before it, so the underlying comparison chain rests entirely on Nvidia's own prior and current claims about its own hardware.

Why it matters

A generational efficiency claim this large changes the economics of running agentic workloads at scale if it holds, which is exactly why the two caveats Nvidia itself included matter more than the headline number: a pending third-party review and an admission that a core agentic capability isn't yet reflected mean the 30x figure is a target the hardware is being marketed toward, not yet a confirmed result. Anyone sizing infrastructure spend against this claim should wait for SemiAnalysis's own published review before treating it as more than Nvidia's stated ambition for its own chip.