HeadlinesBriefing favicon HeadlinesBriefing.com

Qwen3.8 27B Quantization: 4-bit Q4_K_M Holds Up

Hacker News •
×

Benchmarking Qwen3.8 27B quantizations shows the 4-bit Q4_K_M (17 GB) matches full BF16 performance on Terminal-Bench 2.1 and GPQA Diamond, fitting on a 24 GB GPU like RTX 4090 with room for 64k context tokens. The 2-bit UD-Q2_K_XL (10.7 GB) sees slight drops, while 1-bit UD-IQ1_S (6.2 GB) collapses to random chance on GPQA Diamond. Effort level (xhigh default) significantly impacts scores, requiring ~8k reasoning tokens for best results.

IFBench shows no change down to 2-bit. Tests used Unsloth quantizations on Modal GPUs (~$3,000 cost), replicating official Qwen results. Compression hits a cliff at 1-bit, where longer reasoning worsens performance.

Q4_K_M remains the sweet spot for consumer hardware without noticeable quality loss on key benchmarks.