Emergent Object Binding Has a Finite Spatial Horizon
Not provided
Abstract
This paper explores how pretrained Vision Transformers encode object binding through self-supervised pretraining, revealing that binding is a local phenomenon with a finite spatial horizon.
Reality Card
The study demonstrates that the probability of two image patches being recognized as belonging to the same object decreases with distance, following an exponential decay pattern with a finite length scale.
The binding probability falls off monotonically with distance and levels off at a nonzero floor, indicating a local spatial coherence.
The findings are based on a limited number of layers probed in the DINO models, which may affect the generalizability of the results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.