HeadlinesBriefing favicon HeadlinesBriefing.com

NVIDIA's Rubin GPU: A Deep Dive

TechPowerUp News •
×

NVIDIA's new "Rubin" GPU architecture marks a significant leap in scaling computing across data centers. This accelerator boasts up to 224 streaming multiprocessors and 288 GB of HBM4 memory, with 336 billion transistors. It features 896 Tensor Cores and a third-generation Transformer Engine, delivering 50 Peta FLOPS at sparse NVFP4 operations.

NVIDIA organizes resources into Graphics Processor Clusters (GPCs) with a centralized L2 cache, managed by the Giga Thread Engine. The GPU can be partitioned into virtual GPUs using MIG Control. The HBM4 memory, provided by partners, offers up to 22 TB/s peak bandwidth. NVLink 6 enhances communication with 3,600 GB/s fabric bandwidth for GPU-to-GPU and 1,800 GB/s for CPU-GPU via NVLink-C2C.

"Rubin" addresses GPU communication latency with device-initiated transfers. A single NVL72 rack houses 72 GPUs. The "Vera Rubin" NVL72 system uses Intelligent Power Smoothing to reduce average power consumption by ~10% and peak power by ~20% compared to the previous "Grace Blackwell" NVL72, further optimized by NVIDIA DSX Max LPS.