Does the truthfulness signal survive code-mixing? Probing hidden states for hallucination detection in Hinglish
Read the original at arxiv.org→arXiv:2609.22138v1 Announce Type: new Abstract: Hidden-state hallucination probing - training a linear classifier on an LLM's internal activations to detect whether a generated answer is faithful to the input - is...
Original headline: "Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.CL lead source Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish