Curved Inference II: Sleeper Agent Geometry — Extending interpretability beyond probes
Read the original at arxiv.org→arXiv:2608.24037v1 Announce Type: new Abstract: This paper extends Anthropic's Sleeper Agents research [1], which showed artificial backdoors persist through safety training & can be detected by linear probes with...
Original headline: "Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes"
Coverage timeline
- Aug 26, 04:00 UTC arXiv cs.CL lead source Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes