STAR-YOLO26: Synthetically Trained Amodal Recovery of Broccoli Heads
G. Christofi (TU Delft - Electrical Engineering, Mathematics and Computer Science)
C.C.S. Liem – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
Jeroen Wildenbeest – Mentor (Hogeschool Inholland)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Accurate biomass estimation remains critical for robotic harvesting, yet standard computer vision models severely underestimate crop size when broccoli heads are occluded by leaves. This paper investigates whether a high speed, edge viable modal model (YOLO26) can be trained to amodally segment occluded broccoli heads using targeted synthetic data instead of labor intensive manual annotations. To execute this, an automated processing pipeline superimposed leaves over fully visible crop backgrounds to synthesize realistic leaf occlusions. This generated synthetic dataset was subsequently deployed to train the network, establishing the data-driven amodal architecture STAR-YOLO26. The resulting model was evaluated against the heavier, two-stage ORCNN architecture across standard localization metrics, processing speed, and sizing parameters. The results indicate that STAR-YOLO26 successfully infers hidden crop boundaries producing comparable results to ORCNN with the primary differences being that STAR-YOLO26 systematically underestimates while ORCNN overestimates the shape area and that ORCNN performs better under heavier occlusions. Crucially, STAR-YOLO26 delivers these results while providing a massive 700% increase in processing speed compared to ORCNN. In conclusion, modal single-stage models demonstrate the capacity to generate amodal predictions when supplied with targeted training distributions, offering a precise, real-time, and scalable framework for automated crop segmentation and subsequent size estimation.
https://data.4tu.nl/datasets/aaef7e9c-47a5-44d4-af6c-1d2eb41b0f28
This dataset contains synthetically modified images based on the Data underlying the publication: Image-based size estimation of broccoli heads under varying degrees of occlusion. Version 3. 4TU.ResearchData. dataset. https://doi.org/10.4121/13603787.v3 by P.M. (Pieter) Blok; van Henten, Eldert; Frits van Evert; Gert Kootstra, used under CC BY-NC-SA 4.0. The original images were altered by overlapping leaves on fully visible broccoli heads to create varying occlusions (one leaf per image). The annotations included with these synthetically occluded images are the modal annotations of the original (fully visible) broccoli head, serving as the amodal mask for the model to train on. This dataset was used to train STAR-YOLO26, a YOLO26 model capable of amodal recovery of broccoli heads under occlusion, as part of a soon to be published bachelor thesis. In accordance with the ShareAlike clause of the original dataset, this derivative dataset is released under the same CC BY-NC-SA 4.0 terms.