I trained a 3.87B MoE model from scratch with 1.45B active parameters on 86.5B tokens
Read the original at www.reddit.com→First of all, thank you for reading. I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons. Apex-2 - Architecture: Decoder-only MoE,...
Original headline: "I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens"
Coverage timeline
- Oct 4, 15:50 UTC r/LocalLLaMA lead source I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens