Labs should still produce strong non-thinking (instruct) models, as newer models’ instruct mode offers no clear performance gains and may reduce coding bench performance vs older 3.6/3.8 variants.
Read the original at www.reddit.com→The results are in, no thinking / instruct mode for new models degrade performance more than on old models such as 3.6 vs 3.8, where 3.6 takes the lead on several coding benches in instruct mode. I would ask the labs...
Original headline: "Should we plead opensource labs to still produce great non thinking (instruct) models?"
Coverage timeline
- Oct 2, 18:10 UTC r/LocalLLaMA lead source Should we plead opensource labs to still produce great non thinking (instruct) models?