Qwen-2B-RCOL: dynamic low-bit quantization with VRAM-budgeted approximate algorithm and improvements at 1–2 bits
Read the original at www.reddit.com→Hey guys - primarily a research release with working models, Not so much a model as a psuedo-new quantization technique. I've been experimenting with a modification of ISTALab's RCO algorithm that can quantize models...
Original headline: "Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)"
Coverage timeline
- Oct 9, 06:14 UTC r/LocalLLaMA lead source Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)