When keywords drop but classifiers hold: soft refusals under KV cache compression
Read the original at arxiv.org→arXiv:2609.31678v1 Announce Type: new Abstract: KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with...
Original headline: "When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression"
Coverage timeline
- Sep 29, 04:00 UTC arXiv cs.LG lead source When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression