The shelf audit problem you are wrestling with is not really a detection problem, and it is not even a vision problem in the traditional sense. It is an information architecture problem hiding inside a model architecture. You have a YOLO detector that crops perfectly, and then you ask a single embedding to carry the entire burden of distinguishing a 1.25L bottle from a 2L bottle when the only visual difference is a few pixels of text that your preprocessing pipeline literally erases via letterboxing. That is not a failure of your approach; it is a fundamental mismatch between the granularity of the question and the capacity of the representation.
The fact that DINOV2, SigLIP2, and OpenCLIP all fail in the same way should tell you something important: this is not a model quality issue. These are models trained to capture semantic similarity at scale, and for them, a Coke bottle is a Coke bottle whether it is 1.25L or 2L. The text is the only discriminator, and you have already identified that the text is invisible after resizing to 224. So the real question is not which embedder to fine-tune, but whether a single global embedding is the right tool for this job at all. The answer, for your use case, is probably no. You need a hybrid approach: use the embedding for what it is good at, which is narrowing candidates to the same brand and general shape, and then let a second stage, perhaps OCR or a lightweight classifier on the cropped label region, make the final call. This is not giving up on embeddings; it is being honest about their limits.
This connects to a broader pattern we have seen in our community. When Exploring Smarter Paths to Text Clustering With Large Language Models discusses the gap between semantic similarity and exact task requirements, it is the same principle. And when Has anyone measured specification ambiguity as a predictor of correlated failure across model families? probes why different models fail in similar ways, you are seeing that exact phenomenon. Your embedders are not broken; they are all encoding the same ambiguity. The difference is that you have the chance to design around it, whereas those other projects are still searching for the root cause.
What we would tell you, and what we would tell anyone building this kind of tool, is to stop trying to make one model do everything. Your instinct to fine-tune on hard negatives is reasonable, but it will likely be slow and brittle, especially with only a couple of reference photos per SKU. Instead, treat identification as a two-stage process. Stage one is the embedding, which gives you a shortlist of plausible candidates. Stage two is a targeted verification, either by running OCR on the label region or by training a small classifier on the specific text patterns that matter for your SKUs. This is not a compromise; it is a more robust architecture. And it is the kind of practical, human-centered design that actually ships. The takeaway you can quote: **A single embedding is a recall tool, not a precision tool. Build your pipeline to use it as such, and add a second stage for the final decision.** That is how you solve size variants without losing your mind.