Distillation for Incrimination and Distillation for Capabilities
Read the original at arxiv.org→arXiv:2610.11012v1 Announce Type: new Abstract: Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling...
Coverage timeline
- Oct 9, 04:00 UTC arXiv cs.AI lead source Distillation for Incrimination and Distillation for Capabilities