Papers/2609.16233
🧪 Test?View on arXiv

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

Not provided

multimodalreasoning3D understanding
2609.16233
Builder Relevance
80%
1h ago

Abstract

SceneBench introduces a benchmark for evaluating vision-language models in 3D spatial reasoning, addressing limitations in current datasets and evaluation tasks.

Reality Card

Core Claim

SceneBench provides a comprehensive benchmark of 966 photorealistic 3D scenes with hierarchical annotations, revealing significant performance drops in models for complex reasoning tasks.

Method / Result

Models achieve up to 85% accuracy for basic detection tasks but drop to 60% for hierarchical reasoning.

Limitations

The benchmark relies on a human-in-the-loop pipeline, which may affect reproducibility due to the extensive human input required.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers