SLMs and QAT: labs debate training precision versus efficient quantization and the push to smaller models like Nanbeige, Liquid, and Qwen; LFM2.5 2.6B released and planned for Q8 testing
Read the original at old.reddit.com→I know many labs trying to shrink deployment costs and increase efficiency. While I do think that that is fine and dandy, I do sometimes question why they bother going so small, and yet training with all 16 bits....
Original headline: "SLMs & QAT"
Coverage timeline
- Aug 5, 10:38 UTC r/LocalLLaMA lead source SLMs & QAT