Intrinsic Structure: Spectral identifiability for mechanistic interpretability
Read the original at arxiv.org→arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of...
Original headline: "Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability"
Coverage timeline
- Aug 12, 04:00 UTC arXiv cs.LG lead source Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability