Open-Weight Masked Introspection finds eight open-weight models cannot reliably report on whether their internal computation was altered
Read the original at arxiv.org→arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own...
Original headline: "Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation"
Coverage timeline
- Aug 24, 04:00 UTC arXiv cs.AI lead source Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation