Qwen3.8-27b on RTX 3090 achieves 82 tps per single request and up to 672 tps peak with quantization and KV cache optimizations.
Read the original at old.reddit.com→Hi, After a long night of optimizations, I believe I have made the fastest inference engine for Qwen3.6-28B on a 3090. Quick metrics: - 250w power capped - Up to 195k context (ships with 150k for safety though) - 82...
Original headline: "Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak"
Coverage timeline
- Aug 16, 19:38 UTC r/LocalLLaMA lead source Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak