Full 1M context on a single RTX5090 with DDR5 desktop setup using vLLM CPU/RAM offloading; ~800 tps per prompt and 15+ tps decode
Read the original at old.reddit.com→First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:...
Original headline: "[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]"
Coverage timeline
- Aug 4, 14:06 UTC r/LocalLLaMA lead source [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
- Aug 4, 15:00 UTC r/LocalLLaMA Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark