SWORD: Wikidata-based distortions reveal cross-lingual inconsistencies in LLM factual error rejection
Read the original at arxiv.org→arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual...
Original headline: "SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection"
Coverage timeline
- Sep 10, 04:00 UTC arXiv cs.CL lead source SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection