HeadlinesBriefing favicon HeadlinesBriefing.com

NVIDIA Rubin CPX AI Accelerator with HBM4 Memory

TechPowerUp News •
×

NVIDIA has revived its Rubin CPX AI accelerator project, now featuring HBM4 memory instead of the previously planned GDDR7. Well-known supply chain analyst Ming Chi Kuo reported the redesign, noting that the Rubin CPX was previously paused after NVIDIA announced its plans for a specialized accelerator derived from the Rubin GPU family for large-scale agentic AI workloads. The redesigned Rubin CPX features 168 GB of HBM4 memory, an upgrade from the original 128 GB GDDR7 plan.

This GPU SKU is housed in a separate rack with configurations ranging from 64 to 256 standalone CPX GPUs per rack. This GPU is dedicated to prefill workloads, powering racks that manage input context for LLMs and KV cache processing. NVIDIA is separating prefill from decode tasks, with CPX GPUs handling prefill and regular Rubin GPUs managing decode.

A rack tray with eight CPX GPUs can handle 1.34 TB of long-context prefill and the associated KV cache. Since the previous design did not require HBM4 integration, NVIDIA has redesigned the package, likely using TSMC's Co Wo S-S or Co Wo S-L for packaging. The chip delivers 30 Peta FLOPS of NVFP4 compute performance on a monolithic die.

It includes four integrated NVENC and four NVDEC video encoders directly on-chip. For LLMs now containing trillions of parameters, having dedicated racks for prefill while others handle decode allows for much faster token delivery. NVIDIA also recommends maintaining a 1:1 ratio of CPX GPUs to regular GPUs in the decode rack.