🧪 Test?View on arXiv
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
Not provided in the content
multimodalreasoningdiagnostic framework
2608.27461
Builder Relevance
2h ago70%
Abstract
The paper presents SciReC, a model-adaptive multimodal academic dialog benchmark for evaluating relational reasoning in multimodal large language models.
Reality Card
Core Claim
The study reveals that relational reasoning is the primary source of error across all models, with Claude 4.6 achieving the best overall relational score of 73%.
Method / Result
Claude 4.6 achieved the best performance on the overall relational score with 73%.
Limitations
Open-source models achieve their lowest scores on spatial relations, indicating potential limitations in their performance across different relational reasoning tasks.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.