Wiring beats blending: what transfers between transformer sizes — and what doesn’t
Read the original at arxiv.org→arXiv:2608.02829v1 Announce Type: new Abstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversion in the Pythia...
Original headline: "Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't"
Coverage timeline
- Aug 5, 04:00 UTC arXiv cs.LG lead source Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't