Papers/2608.12515
๐Ÿงช Test?View on arXiv

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

Not specified in the provided content

fine-tuningmultimodalreasoning
2608.12515
Builder Relevance
60%
6h ago

Abstract

This paper evaluates vision-language models for classifying egocentric robot images into danger levels, highlighting the limitations in proxemic reasoning and spatial grounding.

Reality Card

Core Claim

Qwen-VL with an advanced prompt significantly improves recall for high-danger cases compared to other models.

Method / Result

Fine-tuning yields only modest overall improvements, but targeted prompting enhances high-danger detection.

Limitations

Current VLMs show limited fine-grained proxemic reasoning and spatial grounding, affecting the reliability of danger classification.

โ† Back to all papers