Agreement is not alignment: divergent moral grounds in human and LLM ethical judgments
Read the original at arxiv.org→arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human...
Original headline: "Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments"
Coverage timeline
- Aug 15, 04:00 UTC arXiv cs.AI lead source Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments