DSpark draft model runs extremely slowly (1-2 t/s) with DeepSeek-V4-Flash on llama-server vs MTP
Read the original at old.reddit.com→Hey everyone, I could use some advice on setting up speculative decoding correctly with llama-server. My Hardware: GPUs: RTX 4090 + RTX 6000 Pro (120GB total VRAM) RAM: 32GB I am currently testing the...
Original headline: "Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?"
Coverage timeline
- Aug 8, 22:28 UTC r/LocalLLaMA lead source Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?