Minimum VRAM required to run DeepSeek-V4-Flash-0731 Q4_K_XL at ~30 t/s
Read the original at old.reddit.com→Hello guys, I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageable. Did someone...
Original headline: "Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?"
Coverage timeline
- Jul 31, 19:37 UTC r/LocalLLaMA lead source Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?
- Aug 1, 08:27 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026
- Aug 1, 08:56 UTC r/LocalLLaMA Deepseek V4 Flash 0731. LM Studio loading only into RAM.
- Aug 1, 16:10 UTC r/LocalLLaMA DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
- Aug 1, 21:22 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5