AMD patch reduces MTP buffer overhead, increasing Qwen 27B context from 64K to 149K
Read the original at old.reddit.com→Available context length with and without the patch: Model: QWEN 27B ROCm stock patched Vulkan stock patched IQ4_XS Pure, single 16GB GPU 19.456 76.032 68,352 78,592 Q6_K_L on 16GB + 12GB 64,256 149,248 68,864...
Original headline: "AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B"
Coverage timeline
- Aug 9, 10:21 UTC r/LocalLLaMA lead source AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B