Analyzing self-harm representations in language models: a cross-architecture study
Read the original at arxiv.org→arXiv:2607.21988v1 Announce Type: new Abstract: Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enable timely...
Original headline: "Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study"
Coverage timeline
- Jul 27, 04:00 UTC arXiv cs.CL lead source Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study