Llama.cpp lacks support for changing thinking amount; user requests on-the-fly mode switching with qwen 3.8 27b referenced
Read the original at old.reddit.com→With qwen 3.8 27b being the long thinker of the year, we could really use support for changing thinking modes on the fly. Last I saw it was being worked on but didn't make a ton of progress. It would be really nice...
Original headline: "We really need llama.cpp to support changing thinking amount"
Coverage timeline
- Aug 18, 12:33 UTC r/LocalLLaMA lead source We really need llama.cpp to support changing thinking amount