Study identifies neurons in a frozen BERT-base-uncased model that support AI-text detection using sparse probing on the RAID benchmark across six generators
Read the original at arxiv.org→arXiv:2609.30287v1 Announce Type: new Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We...
Original headline: "A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.CL lead source A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID