🧪 Test?View on arXiv
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Author1, Author2, Author3, Author4, Author5
multimodalpersonalized safetyinput monitoringvisual dominance
2609.04281
Builder Relevance
7h ago80%
Abstract
This paper addresses the issue of personalized safety in vision-language models (VLMs) deployed in high-stakes settings.
Reality Card
Core Claim
The introduction of PRISM, a lightweight input monitor that predicts when a query requires deferral, achieving 0.978 AUC and dominating the safety-utility Pareto frontier across all tested models.
Method / Result
PRISM achieves 0.978 AUC.
Limitations
The benchmark relies on hidden user profiles, which may complicate reproducibility and generalization of results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.