Qwen3.8-27B inference gains reach on RTX 3090 with 138 tps at default sampling and 942 tps at 64 concurrent, following updates to Dflash2 model
Read the original at old.reddit.com→Edit: Title says 134 tps, it's actually 138 -- keep in mind my 3090 is power limited to 250w. Three days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672...
Original headline: "I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090"
Coverage timeline
- Aug 19, 20:28 UTC r/LocalLLaMA lead source I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090