KLQ: Training-free measured rotation quantization outperforms other training-free rotation-based methods on W4A4KV4-bits; Llama 3.2 1B KLQ-quantized beats SpinQuant and nears ReSpinQuant without GPTQ/LDLQ rounding
Read the original at old.reddit.com→First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about quantization and...
Original headline: "KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding."