SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation
Read the original at arxiv.org→arXiv:2609.13198v1 Announce Type: new Abstract: One of the pivotal recent challenges in neural network interpretability is polysemanticity, where a single neuron is activated by multiple, often unrelated concepts,...
Coverage timeline
- Sep 15, 04:00 UTC arXiv cs.LG lead source SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation