What Does the Encoder Actually Decide? A Controlled Comparison of Vision Backbones on Joint Tree Segmentation and Stereo Depth
Author1, Author2, Author3, Author4, Author5
Abstract
This paper investigates the impact of different vision backbones on joint tree segmentation and stereo depth estimation in robotic pruning tasks.
Reality Card
Convolutional and hybrid encoders outperform transformers in joint semantic segmentation and stereo depth tasks, with a small encoder achieving better performance than much larger models.
The best encoder achieved the highest segmentation mIoU and depth δ1, demonstrating that parameter count does not predict model quality.
The study's findings may not generalize beyond the specific dataset and tasks used, limiting reproducibility in different contexts.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.