Question about quantization versus model size in LLMs
Read the original at old.reddit.com→Sorry if this is asked a lot, but I was wondering if there is any clear winner on the Quantization versus Model Size debate? I can run Qwen3.6 27b at Q8, Laguna at Q6, and the new Deepseek Flash at Q3 bit. I am in...
Original headline: "Question about Quant versus Size."
Coverage timeline
- Aug 3, 00:09 UTC r/LocalLLaMA lead source Question about Quant versus Size.