Papers/2608.24935
🧪 Test?View on arXiv

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

William Bu, Author 2, Author 3, Author 4, Author 5

multimodalcomputer visionroboticsagriculture
2608.24935
Builder Relevance
80%
3h ago

Abstract

This study presents a lightweight multimodal vision-language framework for classifying early-stage apple fruitlet anatomical structures, crucial for robotic thinning and precision orchard operations.

Reality Card

Core Claim

The framework achieved high F1-scores (0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle) using a lightweight model optimized for edge deployment.

Method / Result

Achieved a macro-F1 score of 0.93 with millisecond-level patch inference on NVIDIA Jetson hardware.

Limitations

The dataset is limited to 600 images from specific orchards, which may affect generalizability.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers