🧪 Test?View on arXiv
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
Not provided in the abstract
causal inferencelarge language modelsconfidence estimation
2608.23660
Builder Relevance
3h ago70%
Abstract
This paper evaluates the reliability of large language models (LLMs) in making causal judgments and their confidence levels.
Reality Card
Core Claim
LLMs are better utilized as sources of externally validated soft causal priors rather than as direct evidence of causal structure.
Method / Result
Models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges.
Limitations
Conventional confidence estimates from LLMs are unreliable, leading to substantial overconfidence in incorrect predictions.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.