Converting dense models into sparse Mixture-of-Experts without pretraining from scratch; author experiments on Qwen/Qwen2.5-0 with RTX 4060.
Read the original at www.reddit.com→For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch. The Idea If you can turn a dense model into an MoE that only runs...
Original headline: "Converting dense models into Mixture-of-Experts"
Coverage timeline
- Oct 11, 05:45 UTC r/LocalLLaMA lead source Converting dense models into Mixture-of-Experts