Strategies for capping thinking on ds4 flash 0731
Read the original at old.reddit.com→I like the outputs from this model, but DAMN does it over think. Has anyone found a robust fix for this that isn't just capping output tokens? Anyone working on a 'thinking cap' for it? Some combo of llama params, or...
Coverage timeline
- Aug 3, 02:55 UTC r/LocalLLaMA lead source Strategies for capping thinking on ds4 flash 0731