๐งช Test?View on arXiv
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
Not specified in the provided content
fine-tuningmultimodalreasoning
2608.12515
Builder Relevance
6h ago60%
Abstract
This paper evaluates vision-language models for classifying egocentric robot images into danger levels, highlighting the limitations in proxemic reasoning and spatial grounding.
Reality Card
Core Claim
Qwen-VL with an advanced prompt significantly improves recall for high-danger cases compared to other models.
Method / Result
Fine-tuning yields only modest overall improvements, but targeted prompting enhances high-danger detection.
Limitations
Current VLMs show limited fine-grained proxemic reasoning and spatial grounding, affecting the reliability of danger classification.