Qwen 3.8 Flash-Next on AMD 7900 XTX shows echoing prior reasoning across 13,000 turns due to preserve_thinking: true feeding back into prompts
Read the original at www.reddit.com→I was running Qwen3.8-Flash-Next (base and later Swift-1.5 NVFP4) on a paid pod seat for about a month, and it started answering with short strings that echoed its own earlier reasoning. I assumed the context was...
Original headline: "I ran Qwen 3.8 Flash-Next on my AMD 7900 XTX at 500k context. All local and what a shift it has been."
Coverage timeline
- Oct 11, 21:44 UTC r/LocalLLaMA lead source I ran Qwen 3.8 Flash-Next on my AMD 7900 XTX at 500k context. All local and what a shift it has been.