HeadlinesBriefing favicon HeadlinesBriefing.com

Lossless GLM-5.2 Compression: 25% Memory Reduction

Hacker News •
×

A recent experiment has demonstrated a significant reduction in memory usage for the GLM-5.2 model through lossless compression techniques. A full scan of the model revealed that a K15 charged-format accounting method could reduce memory by 30.168%.

Furthermore, a separate byte-split representation was successfully decoded bit-for-bit across all 59,509 BF16 tensors, achieving a 24.967% reduction in memory. This byte-split method reconstructs the high byte from a codebook, indices, and an escape stream, while the low byte is kept verbatim, ensuring exact fidelity.

While the K15 layout was not independently decoded at the GLM scale, its accounting suggests potential for greater compression. Early runtime evidence from a dense 12-bit prototype showed a speedup of 0.733 times BF16 GEMV time on an A40, though sparse escape correction was not fused into this timing. The researchers highlight that exact lossless speedup, a physical K15 container, and end-to-end serving integration remain open areas for future work. Prior work from Zip NN and DFloat11 established lossless BF16 exponent compression, with Zip Serv being a close prior runtime design. This experiment utilizes a distinct representation, focusing on per-tensor 4-bit codes with sparse exact escapes.