Papers/2609.04276
🧪 Test?View on arXiv

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

Not provided

failure predictioninterpretabilitymultimodalautoencoders
2609.04276
Builder Relevance
70%
7h ago

Abstract

This paper investigates the use of Sparse Autoencoders for interpretable failure prediction in Vision-Language Models.

Reality Card

Core Claim

The proposed framework outperforms existing methods in failure prediction while providing improved interpretability.

Method / Result

The framework shows improved performance in failure prediction compared to evaluated baselines.

Limitations

The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers