Papers/2609.13158
🧪 Test?View on arXiv

TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

Yuqing Yu, Author 2, Author 3, Author 4, Author 5

VQAdocument-level reasoningmultimodalbenchmarking
2609.13158
Builder Relevance
80%
2h ago

Abstract

This paper introduces TestHallVQA, a multi-image VQA benchmark that addresses the limitations of existing benchmarks by incorporating document-level reasoning and contextual redundancy.

Reality Card

Core Claim

TestHallVQA provides a comprehensive benchmark that quantifies LVLMs' reasoning capabilities and robustness against document-level redundancy.

Method / Result

The novel metric F1-R² quantifies LVLMs' computational reasoning capability and evidence retrieval robustness.

Limitations

The impact of irrelevant visual tokens on performance has not been systematically quantified in existing benchmarks.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers