Tech & AI News
Hacker News

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

The 4-bit Q4_K_M quantization of Qwen3.8 27B (17 GB) matches the full BF16 model on the Terminal-Bench 2.1 coding benchmark and fits on a 24 GB GPU while retaining about 64 k tokens of context. In contrast, 1-bit quantization reduces performance to near-random on GPQA Diamond, and lower-bit models (2-bit) show modest degradation on GPQA and IFBench but still retain reasonable scores.