Copying before suppression: a below-chance dip during language model training
Read the original at arxiv.org→arXiv:2610.04119v1 Announce Type: new Abstract: Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task....
Original headline: "Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.CL lead source Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?