Llama.cpp adds GPU-based sampling for MTP, claiming about an 8% speedup in tok/s on Qwen3.6-35B with a 5090; observed ~4% speedup on Nvidia P40 in tests.
Read the original at old.reddit.com→Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and observed a 4%...
Original headline: "Llama.cpp PR 8% speed boost"
Coverage timeline
- Aug 4, 12:16 UTC r/LocalLLaMA lead source Llama.cpp PR 8% speed boost