Safety beyond the interface: detecting harm via latent states in large language models
Read the original at arxiv.org→arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead....
Original headline: "Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models"
Coverage timeline
- Sep 18, 04:00 UTC arXiv cs.AI lead source Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
- Sep 18, 04:00 UTC arXiv cs.CL The Role of Fine-grained Harm Signals in LLM Safety