Cultural misalignment in large language models: evaluation across 63 personas using Wasserstein distance; Qwen3-4B worst on its own population across three countries
Read the original at arxiv.org→arXiv:2609.04485v1 Announce Type: new Abstract: We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63...
Original headline: "Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning"
Coverage timeline
- Sep 7, 04:00 UTC arXiv cs.CL lead source Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning