Qwen 3.8 27B runs slowly on vLLM with RTX 6000 Pro, regardless of thinking effort settings.
Read the original at old.reddit.com→Title pretty much says it all. I’ve deployed Qwen 3.8 27B using vLLM on an RTX 6000 Pro (tried multiple vLLM releases and launch recipes), but I can't get it into a usable state because of crazy long reasoning...
Original headline: "Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking"
Coverage timeline
- Aug 16, 05:47 UTC r/LocalLLaMA lead source Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking