Papers/2609.04281
🧪 Test?View on arXiv

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Author1, Author2, Author3, Author4, Author5

multimodalpersonalized safetyinput monitoringvisual dominance
2609.04281
Builder Relevance
80%
7h ago

Abstract

This paper addresses the issue of personalized safety in vision-language models (VLMs) deployed in high-stakes settings.

Reality Card

Core Claim

The introduction of PRISM, a lightweight input monitor that predicts when a query requires deferral, achieving 0.978 AUC and dominating the safety-utility Pareto frontier across all tested models.

Method / Result

PRISM achieves 0.978 AUC.

Limitations

The benchmark relies on hidden user profiles, which may complicate reproducibility and generalization of results.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers