🧪 Test?View on arXiv
Same Quantity, Different Answer: Numerical Representation Invariance in Language Models
Not specified in the provided content
reasoningevaluationlanguage models
2609.25009
Builder Relevance
1h ago70%
Abstract
The paper investigates how numerically equivalent word problems yield different answers based on their representation in language models.
Reality Card
Core Claim
The study demonstrates that while canonical accuracy is high, the models struggle with orbit correctness and invariance, particularly with scientific notation.
Method / Result
Canonical accuracy ranges from 0.969 to 0.996 across evaluated systems.
Limitations
The main limitation is the collapse of strict-parser performance due to multiplication-form scientific notation being outside the implemented number grammar.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.