The Little Bit Project introduces sub-1-bit large language model compression via latent factorization. Official implementations of Little Bit (NeurIPS 2025) and Little Bit-2 (ICML 2026) are now available on GitHub. Little Bit compresses dense weight matrices into low-rank latent factors, binarizes them, and restores magnitude through learned scales—achieving compression down to 0.1 bits-per-weight while preserving model architecture at inference.
Little Bit-2 improves initialization by aligning latent geometry using Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ), available via the `--use_itq` flag. This opt-in step introduces no inference overhead and enhances Quantization-Aware Training (QAT) compatibility. Supported models include OPT, Llama, Llama 2/3, Phi-4, Qwen2.5, Qwen3, Gemma 2, and Gemma 3.
The codebase requires Python 3.12, CUDA toolkit, PyTorch 2.8.0+cu124, and transformers 4.51.x for reproducibility. Training and evaluation scripts are provided for single- and multi-GPU setups using DeepSpeed.
Source: Hacker News · Summarized by HeadlinesBriefing