September 19, 2026
AI inference rack available now: GB300 NVL72 packs 72 Blackwell Ultra GPUs
As of 19 September 2026, GB300 NVL72 is available through NVIDIA Contact Sales: a single liquid-cooled rack-scale system includes 72 Blackwell Ultra GPUs and 36 Grace CPUs. NVIDIA claims 1.5x greater FP4 Tensor Core FLOPS density and twice the attention performance of Blackwell GPUs. In March 2025, the company planned partner shipments from the second half of 2025.

In March 2025, NVIDIA said GB300 NVL72 would deliver 1.5x higher AI performance than GB200 NVL72. The system is now available, and for AI factories the company cites up to 50x higher output performance than Hopper platforms.
Memory. The product page lists 37 TB of fast memory, 20 TB of GPU memory with bandwidth of up to 576 TB/s, and 17 TB of LPDDR5X CPU memory. In the technical table dated 19 March, fast memory is listed as 40 TB, HBM as 20 TB, and network bandwidth as 14.4 TB/s. NVLink transfers data at up to 130 TB/s.
For inference. NVIDIA offers the open-source Dynamo: it distributes workloads across GPU nodes and separates prefill from decode. Each GPU receives up to 800 Gb/s through ConnectX-8 SuperNIC, while the platform works with Quantum-X800 InfiniBand and Spectrum-X Ethernet.
What comes next for GB300 NVL72 in infrastructure will be determined by real-world inference benchmarks after partner deliveries.
