🧪 Test?View on arXiv
TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams
Yuqing Yu, Author 2, Author 3, Author 4, Author 5
VQAdocument-level reasoningmultimodalbenchmarking
2609.13158
Builder Relevance
2h ago80%
Abstract
This paper introduces TestHallVQA, a multi-image VQA benchmark that addresses the limitations of existing benchmarks by incorporating document-level reasoning and contextual redundancy.
Reality Card
Core Claim
TestHallVQA provides a comprehensive benchmark that quantifies LVLMs' reasoning capabilities and robustness against document-level redundancy.
Method / Result
The novel metric F1-R² quantifies LVLMs' computational reasoning capability and evidence retrieval robustness.
Limitations
The impact of irrelevant visual tokens on performance has not been systematically quantified in existing benchmarks.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.