Papers/2608.12368
๐Ÿงช Test?View on arXiv

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

Not provided

ethical AIalignmentmoral reasoning
2608.12368
Builder Relevance
70%
6h ago

Abstract

The paper argues that agreement in judgments between humans and large language models does not imply alignment in moral reasoning.

Reality Card

Core Claim

The study demonstrates that while LLMs may agree with human judgments, they often rely on different moral grounds, highlighting the need for deeper analysis beyond label agreement.

Method / Result

The analysis involved a curated 500-item ETHICS-derived benchmark, revealing systematic divergence in moral grounds between human annotators and LLMs.

Limitations

The study's findings may not be generalizable across all LLMs or moral domains due to the specific benchmark used.

โ† Back to all papers