Muse-Glimmer reasoning traces differ noticeably from qwen and gemma models, raising questions about their behavior
Read the original at old.reddit.com→Just downloaded the model, UD-Q5_K_XL quant, asked it to generate a long story to test out reasoning and speed with dflash (super fast btw, ~ 90 to 160 tok/s on a 5090 depending on task) and was surprised at the...
Original headline: "Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys"
Coverage timeline
- Aug 11, 00:38 UTC r/LocalLLaMA lead source Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys
- Aug 11, 03:25 UTC r/LocalLLaMA 1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases