10% faster decode with Q4_K MTP draft model using Gemma 4 31b
Read the original at old.reddit.com→(Disclaimer: I am a noob and don’t know what I am doing) Gemma 4 31b unsloth/gemma-4-31B-it-qat-GGUF I took the f16 MTP draft model and quantised it to Q4_K (instead of Q4_0 of unsloth) and gained around 10% in...
Original headline: "10% faster decode with Q4_K MTP draft model with Gemma 4 31b"
Coverage timeline
- Aug 6, 16:43 UTC r/LocalLLaMA lead source 10% faster decode with Q4_K MTP draft model with Gemma 4 31b