Toxicity priors show conditional reliability in multilingual and code-mixed abuse detection; English toxicity, Indic abuse, and rule-based severity cues help only in certain linguistic and abuse contexts (ToxGate).
Read the original at arxiv.org→arXiv:2607.15861v1 Announce Type: new Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We...
Original headline: "Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection"