Interaction-aware circuit discovery in language models reveals context-dependent component interactions; arXiv:2610.04017v1
Read the original at arxiv.org→arXiv:2610.04017v1 Announce Type: new Abstract: Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses...
Original headline: "Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.LG lead source Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models