SharedSAE uses a single feature dictionary across language models to replace per-model sparse autoencoders
Read the original at arxiv.org→arXiv:2609.04344v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here,...
Original headline: "SharedSAE: One Feature Dictionary Across Language Models"
Coverage timeline
- Sep 7, 04:00 UTC arXiv cs.LG lead source SharedSAE: One Feature Dictionary Across Language Models