🧪 Test?View on arXiv
CARAT: Do Materials LLMs Reason or Recite?
Not provided in the abstract
reasoningfine-tuningevidence injection
2609.38340
Builder Relevance
1h ago70%
Abstract
The paper investigates whether materials LLMs reason from crystal structures or simply recite pre-existing answers.
Reality Card
Core Claim
The grounded view in CARAT improves accuracy by 17.3 points over formula inputs on the benchmark's hardest families.
Method / Result
GraphSpace outperforms a plain periodic graph by 19.3 points.
Limitations
The benchmark's design may favor certain types of evidence presentation, affecting reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.