October 1, 2026
NVIDIA Vera Rubin's gains depend on the workload: 2x in SemiAnalysis testing
On Sep 14, SemiAnalysis measured a 2.09-fold advantage for Vera Rubin over GB300 at the same response speed.

NVIDIA estimates the cost of an AI factory at $60 million per megawatt of power capacity, according to figures from Oct 1. The company links the return on investment to performance, equipment lifespan and the ability to keep it busy with different workloads.
Conditions change the results. NVIDIA claims Vera Rubin NVL72 delivers more than 30 times the performance per megawatt of GB300 NVL72 on DeepSeek V4 Pro. The company estimates that the cost per million tokens falls by a factor of up to 45. SemiAnalysis compares the systems at a fixed generation speed for the user, and the result also depends on the software serving the requests.
| Measurement | Vera Rubin | GB300 | Rubin advantage | | --- | --- | --- | --- | | SemiAnalysis, Sep 14, 2026: 100 tokens/s for the user | 59.4 million tokens/s/MW | 28.5 million with Dynamo SGLang | 2.09 times | | SemiAnalysis, Sep 14, 2026: 170 tokens/s for the user | Comparison with GB300 TRTLLM | TRTLLM | 62.9 times | | SemiAnalysis, Sep 14, 2026: 170 tokens/s for the user | Comparison with GB300 SGLang | SGLang | 5.56 times |
Benchmarking on your own server. SemiAnalysis provides instructions in the InferenceX documentation for running AgentX-Harness against an OpenAI-compatible server. Running it requires Git, uv, Python 3.11 and aiperf. The example generates load for 3,600 seconds, or one hour.
Older GPUs continue to generate revenue beyond their estimated useful life. According to NVIDIA's Oct 1 figures, Microsoft operated V100 GPUs for 8.4 years against an accounting useful life of 6 years. CoreWeave extended reservations for A100 GPUs, introduced in 2020, through 2029.
Operators use several GPU generations simultaneously and distribute workloads across them. NVIDIA reports that Pinterest uses 14,000 GPUs from the Blackwell, Hopper and earlier generations to fine-tune and run a model that processes images and text. Texas A&M University keeps its supercomputer at 95–98% utilization through molecular simulations and AI-assisted drug discovery: it runs 26 projects from 7 organizations.
To design this infrastructure, NVIDIA offers the Enterprise AI Factory Design Guide, covering hardware and software architecture, AI agent management and validated deployment configurations.
