🧪 Test?View on arXiv
Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
multimodalinterpretabilitymodel failures
2610.00024
Builder Relevance
1h ago60%
Abstract
The paper investigates mid-layer interpretability in vision-language models, revealing that while mid layers encode ground-truth answers in a significant percentage of errors, this information does not influence final predictions.
Reality Card
Core Claim
The study demonstrates that mid-layer information in vision-language models is not causally active for final predictions, despite encoding ground-truth answers in a high percentage of errors.
Method / Result
Residual-stream patching yields 0% non-trivial flip at the layer level across all three architectures tested.
Limitations
The findings are based on oracle labels, which may not be deployable in practical applications.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.