Papers/2608.23660
🧪 Test?View on arXiv

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

Not provided in the abstract

causal inferencelarge language modelsconfidence estimation
2608.23660
Builder Relevance
70%
3h ago

Abstract

This paper evaluates the reliability of large language models (LLMs) in making causal judgments and their confidence levels.

Reality Card

Core Claim

LLMs are better utilized as sources of externally validated soft causal priors rather than as direct evidence of causal structure.

Method / Result

Models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges.

Limitations

Conventional confidence estimates from LLMs are unreliable, leading to substantial overconfidence in incorrect predictions.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers