DFlash2 speeds Qwen 3.8 27B up to 4 times on decoding throughput, according to tests comparing baseline, MTP, DFlash, and DFlash2.
Read the original at old.reddit.com→llama.cpp pr #27342 adds dflash2, so i rented an rtx 6000 and ran the same four prompts through four decoding setups on qwen3.8 27B median results over the four tasks: baseline 47.4 tok/s mtp 114.7 tok/s dflash 99.3...
Original headline: "DFlash2 speeds Qwen 3.8 27B up to 4 times"
Coverage timeline
- Aug 19, 18:10 UTC r/LocalLLaMA lead source DFlash2 speeds Qwen 3.8 27B up to 4 times