Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens; how middle layers break and how 4 anchor blocks fixed it
Read the original at www.reddit.com→Standard distillation usually means burning weeks of compute and billions of tokens hoping the student model eventually mimics the teacher. We wanted to see what happens if you skip backprop entirely and treat...
Original headline: "Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)"
Coverage timeline
- Oct 6, 08:15 UTC r/LocalLLaMA lead source Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)
- Oct 6, 09:41 UTC r/LocalLLaMA A 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)