Qwen3.8-27B runs with FP8 on dual GPUs (RTX 5090 and RTX 4070 Ti Super) with 262K context and vLLM 0.30.0 benchmarks show 2k prefills, 2,922 t/s decode (code) at 0.7 s; 32k depth.
Read the original at www.reddit.com→Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB, PCIe x1) = 48 GB VRAM 32 GB DDR5-6200, Arch Linux, KDE on the 5090 Setup Huihui Qwen3.8-27B abliterated INT8 W8A16 + DFlash2 drafter (K=7), vLLM 0.30.0,...
Original headline: "Qwen3.8-27B FP8 dual GPUs"
Coverage timeline
- Sep 27, 11:15 UTC r/LocalLLaMA lead source Qwen3.8-27B FP8 dual GPUs