Representational simplicity and circuit size dissociate in a threshold-dependent way; a controlled test via adversarial training
Read the original at arxiv.org→arXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer....
Original headline: "Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training"
Coverage timeline
- Sep 30, 04:00 UTC arXiv cs.AI lead source Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training