🧪 Test?View on arXiv
Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction
Author1, Author2, Author3, Author4, Author5
affordance predictionpart groundingmultimodalmodel evaluation
2609.13225
Builder Relevance
2h ago80%
Abstract
This paper identifies part grounding as the primary bottleneck in vision-language models' ability to predict affordances for manipulation tasks.
Reality Card
Core Claim
Naming the target part significantly improves action accuracy and recall in affordance prediction tasks across multiple models.
Method / Result
Naming the target part raised action accuracy by 0.32 to 0.63 across all models tested.
Limitations
The study documented measurement errors that could affect the reproducibility of results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.