Qwen3.8-27b outperforms larger models on benchmarks, but results are based on bf16 weights; practical runs use 4-bit quantization due to hardware limits, meaning the downloaded artifact differs from the measured one.
Read the original at old.reddit.com→qwen3.8-27b looks genuinely impressive on the benchmark tables - beating models many times its size on some of them. but those numbers come from bf16 weights, and nobody here is running a 27b at bf16. we're running...
Original headline: "we benchmark models nobody actually runs"
Coverage timeline
- Aug 17, 21:53 UTC r/LocalLLaMA lead source we benchmark models nobody actually runs