SDM-D: The Label Factory for Fruit Vision
How a segment-then-prompt pipeline uses SAM2 and OpenCLIP to turn unlabeled orchard images into training data for lightweight detectors.
- SDM-D treats foundation models as a labeling engine first and a detector second, which is why it matters in agriculture.
- Its real trick is splitting geometry from semantics, using SAM2 to find masks and OpenCLIP to decide what they are.
- The repo turns pseudo-labels into standard training data, then distills that supervision into lean models that can run in the field.
- Its value is not one-shot detection quality alone, but a faster path from unlabeled orchard images to deployable data and models.
The hard part in fruit vision is not always inference. It is annotation. Drawing masks for strawberries, leaves, stems, and clutter is slow, expensive, and brittle when lighting changes or fruit overlaps.
SDM-D attacks that bottleneck directly. It uses SAM2 to propose boundaries, OpenCLIP to decide what each region means, and then distills the result into training data for smaller detectors.
The real bottleneck is not inference. It is labeling.
That framing changes the whole project. SDM-D is not trying to beat every detector on the first pass. It is trying to manufacture the supervision that good detectors usually depend on.
In precision agriculture, that matters more than a flashy demo. Orchards are crowded, occluded, and full of lookalikes. Manual labeling does not scale well when every image contains dozens of partial fruit and tangled stems.
What SDM-D actually does
The name gives the system away. SDM-D stands for Segmentation-Description-Matching-Distilling. That is the whole idea in one chain.
First, SAM2 finds object boundaries without caring whether the object is fruit, leaf, or stem. Then OpenCLIP compares each region against a prompt bank. Finally, the accepted labels are exported into standard training formats so a smaller model can learn from them.
# Conceptual flow, simplified
masks = sam2.generate(image)
labels = []
for mask in masks:
region = crop(image, mask)
label = openclip.match(region, prompts)
labels.append((mask, label))
export_to_training_format(labels)
train_small_detector(labels)
Why “segment-then-prompt” is the key move
This is the conceptual center of the repo. Most open-vocabulary systems try to solve geometry and semantics at the same time. SDM-D splits them apart.
That split is powerful because the two problems fail differently. Segmentation is about where the object ends. Prompt matching is about what it is. In orchards, where objects overlap and look alike, decoupling the two reduces the chance that a class name gets dragged down by bad box geometry, or that a shaky boundary corrupts the label.
The mid-article illustration matters because the project is not a generic detector. It is a labeler that happens to use foundation models. That distinction is the whole editorial trick.
| Workflow | Input effort | Output | Speed / compute burden | Best use case |
|---|---|---|---|---|
| Manual labeling | High human effort per image | Ground-truth annotations | Slow but precise | Small datasets where experts can annotate carefully |
| Grounded-SAM style pipeline | Prompting plus heavy inference | Region proposals and labels | Very heavy compute and VRAM | General open-vocabulary exploration |
| SDM-D | A few descriptive prompts and images | Pseudo-labels plus training data | Heavy once, then distilled | Building a narrow-domain dataset for fruit vision |
| Distilled YOLO deployment | Training on pseudo-labels | Fast edge-ready detector | Low compute at inference | Field deployment on robots or embedded hardware |
Why the prompts matter as much as the model
The prompts are where agricultural knowledge enters the system. A class label like “strawberry” is useful, but it is blunt. A description like “a red strawberry with numerous points, ripe” carries more domain detail and gives the matcher something better to work with.
That means the repo is not only exploiting foundation models. It is letting a human expert encode field knowledge once, then reuse it across a large image set. In practice, that is how you turn a general vision stack into a crop-specific label factory.
These pseudo-labels generated by SDM can serve as supervision for small, edge-deployable models (students), bypassing the need for costly manual annotation. The SDM-D is highly versatile and model-agnostic, with no restrictions on the choice of the student model.
a red strawberry with numerous points, ripe
a green veined strawberry leaf, leaf
a thin brown stem segment, stem
How the repo is wired
The codebase is lean and easy to read at the right altitude. SDM.py orchestrates the pipeline. utils.py handles matching and visualization. The seg2label/ scripts convert masks into detector-friendly formats.
The structure mirrors the idea. One file for orchestration, one for glue, one folder for conversion. That is a good sign in a research repo because the code follows the conceptual split instead of hiding it.
# High-level repo layout
SDM.py # pipeline entry point
utils.py # matching, visualization, helpers
seg2label/ # convert masks to training formats
description/ # text prompts for matching
notebook/ # exploration and tuning
The implementation also hints at seriousness. The orchestration layer uses GPU-aware settings like autocast and tf32, which tells you this is designed for heavy foundation-model inference, not a toy notebook.
From pseudo-masks to deployable detectors
The last step is the one that makes SDM-D more than an academic pipeline. The pseudo-labels are not the endpoint. They become training data for compact models that can run in the field.
That bridge matters because orchard robotics cannot afford to carry a giant foundation model everywhere. SDM-D uses the large model to do the expensive thinking once, then hands the result to a smaller student model for real-world use.
| Stage | What happens | Why it matters |
|---|---|---|
| Foundation model pass | SAM2 and OpenCLIP generate pseudo-labels | Creates supervision without manual annotation |
| Conversion pass | seg2label rewrites masks into standard training data | Makes the labels usable by common detection pipelines |
| Student training | A compact detector learns from the pseudo-labels | Turns research output into a practical field model |
| Deployment | The distilled detector runs on edge hardware | Keeps inference fast and feasible in orchards |
How SDM-D compares with the rest of the field
Compared with manual labeling, SDM-D is less labor-intensive. Compared with Grounded-SAM, it is more focused on creating reusable supervision rather than just delivering a heavy open-vocabulary pass. Compared with a plain YOLO workflow, it gives you the labels you would otherwise have to pay for by hand.
That is the important editorial point: SDM-D is not trying to be the best end-user detector. It is trying to be the best way to produce the detector you actually want.
| Workflow | What it gives you | Main trade-off | Where it shines |
|---|---|---|---|
| Manual labeling | Ground truth | Expensive and slow | Very small datasets |
| Grounded-SAM | Heavy open-vocabulary regioning | High VRAM and latency | General exploration |
| YOLO-World | Direct open-vocabulary detection | Less tailored to one crop | Broad category coverage |
| SDM-D | Pseudo-label factory plus distilled detector | More pipeline complexity | Narrow agricultural domains with scarce labels |
What this repo suggests about the next wave of agricultural AI
SDM-D points to a broader pattern. In niche physical domains, the bottleneck is often not model intelligence. It is dataset creation. The winning move is to use foundation models as a temporary labor force, then compress their output into something smaller, cheaper, and easier to deploy.
That makes SDM-D a blueprint, not just a repo. It shows how to bootstrap a crop-specific dataset from a few good prompts and a lot of unlabeled imagery, then distill that work into a practical system for the field.