🧪 Test?View on arXiv
A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
William Bu, Author 2, Author 3, Author 4, Author 5
multimodalcomputer visionroboticsagriculture
2608.24935
Builder Relevance
3h ago80%
Abstract
This study presents a lightweight multimodal vision-language framework for classifying early-stage apple fruitlet anatomical structures, crucial for robotic thinning and precision orchard operations.
Reality Card
Core Claim
The framework achieved high F1-scores (0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle) using a lightweight model optimized for edge deployment.
Method / Result
Achieved a macro-F1 score of 0.93 with millisecond-level patch inference on NVIDIA Jetson hardware.
Limitations
The dataset is limited to 600 images from specific orchards, which may affect generalizability.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.