HeadlinesBriefing favicon HeadlinesBriefing.com

NVIDIA's 'Rubin' Architecture Details Revealed

TechPowerUp News •
×

NVIDIA has unveiled technical details for its next-generation "Rubin" GPU, designed for "agentic AI" workloads. This architecture aims for significant improvements in inference throughput and energy efficiency, claiming up to 10x more agentic throughput per unit of energy compared to the previous "Blackwell" generation. The "Rubin" GPU features a chiplet design, utilizing TSMC's CoWoS-L packaging to accommodate two compute dies. Each die is at the reticle limit, contributing to a total of 336 billion transistors. The GPU is equipped with up to 224 SMs, 896 Tensor Cores, and 288 GB of HBM4 memory, capable of delivering up to 50 PetaFLOPS for inference and 35 PetaFLOPS for training using NVIDIA's proprietary NVFP4 format.

Key advancements include enhanced Tensor Cores with a third-generation Transformer Engine, optimized for inference bottlenecks and attention mechanisms. The HBM4 memory provides up to 22 TB/s of peak bandwidth, a substantial increase over previous generations. NVIDIA has also improved scale-up communication with "counted writes" and NVLink 6, enabling more efficient data exchange between GPUs in a rack. This system-level optimization prioritizes overall performance, moving beyond single-GPU metrics, and is crucial for the continuous, reasoning-intensive nature of "agentic AI."