๐งช Test?View on arXiv
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
Not provided
ethical AIalignmentmoral reasoning
2608.12368
Builder Relevance
6h ago70%
Abstract
The paper argues that agreement in judgments between humans and large language models does not imply alignment in moral reasoning.
Reality Card
Core Claim
The study demonstrates that while LLMs may agree with human judgments, they often rely on different moral grounds, highlighting the need for deeper analysis beyond label agreement.
Method / Result
The analysis involved a curated 500-item ETHICS-derived benchmark, revealing systematic divergence in moral grounds between human annotators and LLMs.
Limitations
The study's findings may not be generalizable across all LLMs or moral domains due to the specific benchmark used.