HeadlinesBriefing favicon HeadlinesBriefing.com

Micron Explores Near-GPU NAND Flash for LLM Inference

TechPowerUp News •
×

Micron is reportedly exploring high-endurance NAND Flash modules positioned closer to the GPU, a concept termed "near-GPU NAND." This architecture aims to bridge the gap between GPU memory like HBM and traditional storage pools, offering a middle ground in density and durability. By placing hundreds of gigabytes of storage directly on the GPU package or PCB, GPUs could access a fast buffer tier with improved I/O speeds, bandwidth, and read times compared to standard TLC or QLC NAND.

This approach targets massive LLM inference workloads, potentially allowing models to run on systems with fewer GPUs by compensating with additional memory tiers. The near-GPU NAND would be significantly cheaper than HBM or DRAM and sized to system specifications. Micron is not alone; SK hynix and Sandisk have already standardized High-Bandwidth Flash (HBF) to extend GPU HBM space with more durable NAND. All companies face the challenge of overcoming bandwidth and durability gaps to make this a viable complement to standard HBM memory.