Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Help: Project l'm building a shelf audit tool. A photo goes through YOLO, which crops each product, and then I embed the crop and search a small gallery of reference photos to get the SKU. New products should be addable by just dropping in photos, no detector retrain. Detection is basically fine. ldentification is not. Same brand, same bottle, different flavor or size (like 1.25 L vs 2 L) and the nearest neighbor is often the wrong SKU. Correct and wrong scores overlap, so a threshold either misses real products or accepts the wrong one. I tried DINOV2, SigLIP2 and OpenCLIP. Same story. Crops get letterboxed to 224, so the tiny "1.25L" / "2L" text basically disappears. Most SKUS only have a couple of shelf photos as references, not clean studio shots. Has anyone actually shipped something like this? Did you fine-tune the embedder on hard negatives, add OCR as a second check, or give up on one global embedding? Curious what worked for size variants. thanks in advance
[link] [comments]
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.