I pre-trained a 700m model on 18B tokens optimized for Python and Wikitext; TheOneWhoWill/Shibai-700M-Base on Hugging Face
Read the original at old.reddit.com→I know this is the 1000000th new sub billion parameter model out there and probably isn't as good as Qwen 3 0.6B or Qwen 3.5 0.8B but it still packs a decent punch. My intention to to continuously pre-train this...
Original headline: "I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face"
Coverage timeline
- Jul 29, 19:29 UTC r/LocalLLaMA lead source I pre-trained a 700m on 18B tokens optimized for Python and Wikitext | TheOneWhoWill/Shibai-700M-Base · Hugging Face