Papers/2608.14558
🧪 Test?View on arXiv

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

multimodalreasoningcognitivebenchmark
2608.14558
Builder Relevance
70%
1h ago

Abstract

This paper introduces The Unwritten Benchmark, a challenge for multimodal models to perform abstract perceptual reasoning through acousto-kinematic word inference.

Reality Card

Core Claim

Multimodal models like GPT-4o and Gemini 2.5-Pro struggle significantly with acousto-kinematic word inference, achieving less than 10% accuracy compared to over 80% for humans.

Method / Result

Human participants achieve over 80% ordered letter accuracy, while leading models fail to surpass 10%.

Limitations

The models exhibit a paradoxical fusion effect where combining modalities degrades performance, indicating a breakdown in synthesizing perceptual cues.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers