$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork achieves Qwen3.8-Flash-Next at 60–100 t/s decode and 3000+ t/s prefill
Read the original at www.reddit.com→Post title is slightly misleading, I don't think you can get these for $350 each anymore but they're still pretty cheap all things considered. They're Radeon Pro V620's which are older RDNA2 enterprise cloud gaming...
Original headline: "$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill"
Coverage timeline
- Oct 8, 17:12 UTC r/LocalLLaMA lead source $2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill