Papers/2610.00024
🧪 Test?View on arXiv

Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

multimodalinterpretabilitymodel failures
2610.00024
Builder Relevance
60%
1h ago

Abstract

The paper investigates mid-layer interpretability in vision-language models, revealing that while mid layers encode ground-truth answers in a significant percentage of errors, this information does not influence final predictions.

Reality Card

Core Claim

The study demonstrates that mid-layer information in vision-language models is not causally active for final predictions, despite encoding ground-truth answers in a high percentage of errors.

Method / Result

Residual-stream patching yields 0% non-trivial flip at the layer level across all three architectures tested.

Limitations

The findings are based on oracle labels, which may not be deployable in practical applications.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers