27B Q8 and 35B Q6 tested on a 32 GB GPU; neither model proved reliable enough to serve as its own final checker
Read the original at old.reddit.com→I started this because of the recurring question around Qwen 27B vs 35B on a single 32 GB card. In particular, some people reported that the dense 27B model seemed to catch coding errors better than the 35B MoE,...
Original headline: "I tested whether 27B Q8 or 35B Q6 is the better coding model on a 32 GB GPU. The more interesting result: neither was reliable enough to be its own final checker."
Coverage timeline
- Aug 12, 01:01 UTC r/LocalLLaMA lead source I tested whether 27B Q8 or 35B Q6 is the better coding model on a 32 GB GPU. The more interesting result: neither was reliable enough to be its own final checker.