Qwen3.5-9B MTP GGUF shows 35.8% acceptance on CUDA and 91–92% on Vulkan; two Radeon/RADV systems are 77% and 128% faster on Vulkan.
Read the original at old.reddit.com→A new llama.cpp issue (#26750) reports a strange MTP reversal with Qwen3.5-9B: the same Q4_K_M GGUF and b10290 tag made decoding 32% slower than baseline on an RTX PRO 4000 Blackwell/CUDA, while two Radeon/RADV...
Original headline: "Same Qwen3.5 MTP GGUF: 35.8% acceptance/−32% on RTX PRO 4000 CUDA, ~92%/+77–128% on two Radeons"
Coverage timeline
- Aug 8, 11:29 UTC r/LocalLLaMA lead source Same Qwen3.5 MTP GGUF: 35.8% acceptance/−32% on RTX PRO 4000 CUDA, ~92%/+77–128% on two Radeons