DeepSeek releases V4-Flash 0731 with enhanced agentic capabilities, 304B parameters (167GB on Hugging Face) and pricing at $0.14 per million input / $0.27 per million output; Artificial Analysis ranks it ahead of MiniMax M3.
Read the original at simonwillison.net→deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch well...
Original headline: "deepseek-ai/DeepSeek-V4-Flash-0731"
Coverage timeline
- Jul 30, 09:32 UTC r/LocalLLaMA What is the best intelligence/stable model currently for a single GB10/DGX spark?
- Jul 31, 06:04 UTC r/LocalLLaMA DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"
- Jul 31, 06:18 UTC r/LocalLLaMA DeepSeek v4 Flash has a nice bump in Capability
- Jul 31, 06:42 UTC r/LocalLLaMA I predict DeepSeek V4 Flash 0731's Artificial Analysis score to be 57 ± 1 point (Kimi K3 Level)
- Jul 31, 07:41 UTC r/LocalLLaMA New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna
- Jul 31, 07:59 UTC Hacker News (AI) DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
- Jul 31, 08:21 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks
- Jul 31, 12:12 UTC r/LocalLLaMA deepseek-ai/DeepSeek-V4-Flash-0731 on Huggingface
- Jul 31, 12:17 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 Open weight!
- Jul 31, 12:21 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 - GGUF also will come as I can see..
- Jul 31, 15:05 UTC r/LocalLLaMA Deepseek V4 Flash on SlopCodeBench
- Jul 31, 15:50 UTC r/LocalLLaMA Deepseek v4 flash MXFP4 (original quality) ggufs
- Jul 31, 16:52 UTC r/LocalLLaMA Some deepseek-v4-flash 20260731 opinion review
- Jul 31, 17:14 UTC r/LocalLLaMA DeepSeek V4 Flash GA ranks the same as Sonnet 5 and Grok 4.5 on DeepSWE
- Jul 31, 17:30 UTC r/LocalLLaMA DeepSeek V4 Flash unsloth quants are out!
- Jul 31, 17:35 UTC r/LocalLLaMA Initial testing of DeepSeek v4 Flash shows significant improvements in UI/UX design capabilities (despite being token hungry)
- Jul 31, 19:37 UTC r/LocalLLaMA Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?
- Jul 31, 20:29 UTC Hacker News (AI) DeepSeek-V4-Flash-0731 model weights (MIT)
- Jul 31, 21:32 UTC r/LocalLLaMA Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper
- Jul 31, 23:31 UTC r/LocalLLaMA DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head
- Jul 31, 23:59 UTC Simon Willison lead source deepseek-ai/DeepSeek-V4-Flash-0731
- Aug 1, 02:35 UTC r/LocalLLaMA What speeds are everyone getting with deepseek v4 flash 0731?
- Aug 1, 08:27 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731: Models you can run locally now have the intelligence score of the top frontier model from March 2026
- Aug 1, 08:56 UTC r/LocalLLaMA Deepseek V4 Flash 0731. LM Studio loading only into RAM.
- Aug 1, 09:44 UTC r/LocalLLaMA New DeepSeek V4 Flash 0731 vs ChatGPT Luna comparison
- Aug 1, 11:41 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??
- Aug 1, 15:01 UTC r/LocalLLaMA DSv4 Flash 0731 Running on Unoptimized Single 3090 System
- Aug 1, 16:10 UTC r/LocalLLaMA DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
- Aug 1, 21:22 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5
- Aug 1, 21:27 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731 on Bosgame M5 with RTX PRO 6000 Max-Q eGPU
- Aug 2, 01:49 UTC r/LocalLLaMA Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- Aug 2, 07:28 UTC r/LocalLLaMA DeepSeek-V4-Flash 284B on 5.3GB of memory
- Aug 2, 11:22 UTC r/LocalLLaMA NEW Deepseek V4 Flash : MMLU-Pro , GPQA Diamond and truthfulQA ?
- Aug 2, 15:43 UTC r/LocalLLaMA Deepseek-V4-Flash-0731 Dwarfstar on Mac
- Aug 2, 16:13 UTC r/LocalLLaMA Deepseek v4 flash - 100-150 faster t/s in prefill/pp.
- Aug 2, 19:16 UTC r/LocalLLaMA DeepSeek-V4-Flash-0731: When Low is higher than High
- Aug 2, 20:27 UTC r/LocalLLaMA DeepSeek V4 @ IQ3XXS on M1 Ultra 128GB- 16 tok/s in LM Studio after patch
- Aug 3, 01:06 UTC r/LocalLLaMA Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8!
- Aug 3, 07:21 UTC r/LocalLLaMA Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090
- Aug 3, 15:19 UTC r/LocalLLaMA DeepSeek V4 Flash 0731 - Happy Numbers (700pp/18tg) and Thoughts
- Aug 3, 16:04 UTC r/LocalLLaMA I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!
- Aug 3, 19:20 UTC r/LocalLLaMA Speculative decoding with deepseek v4 flash 0731?
- Aug 3, 20:25 UTC r/LocalLLaMA DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config