There is a particular kind of confidence that comes from building something small and honest, and that is exactly what the author of peekaboolean has delivered. While the broader field chases ever-larger generative models that produce fluent prose about an image, this project asks a quieter, more practical question: what if you don't need a paragraph, just a probability? The result is a 500M-parameter vision-language model that answers typed questions about an image in roughly 400 milliseconds on a laptop, returning calibrated probabilities for yes/no, choice, and rubric-based scoring tasks. No generated text, no parsing ambiguity, no hallucinated adjectives. Just the answer, scoped to the options you provided. That is a refreshingly direct approach, and it works because it refuses to treat the problem as a language generation task when it is really a classification task wearing a trench coat.
The clever core is that the model never writes a sentence. Each option becomes its own prompt, scored by the difference between the logits for "Yes" and "No" from the backbone's own language model head. This means the model is not free to drift into an unprompted explanation; it either agrees or disagrees with the proposed answer, and a softmax over the options gives you a distribution. The authors also fit a temperature per question type, option count, and image size, which is the kind of calibration detail that separates a demo from a tool. They also made a smart choice in using a 30B teacher (Qwen3-VL) to generate and label roughly 175,000 questions across 62,000 images, then trained the small student to mimic the teacher's probabilities rather than its text. That distillation, combined with a rejection of the teacher's own blind spots, like answering a question the same way with a blank image, shows a discipline that most research projects lack. The numbers on the held-out split are strong: 0.94 balanced accuracy on yes/no, 0.78 accuracy on choice, and 0.87 on DocVQA, which is remarkable for a model that runs in under half a second on an M1 Pro.
But here is where the editorial take sharpens. The author is honest about a critical limitation: the "teacher" rows measure how closely the student copies a 30B model, not how well it performs against human judgment. No human has checked those labels. That is not a small footnote; it is the difference between a useful engineering artifact and a reliable product. The public VQA data alone taught a "benchmark dialect," and the real-format requests only improved after teacher-written data was added. So the model is good at mimicking the teacher's style of judgment, but we still do not know if that judgment is correct for the messy, context-rich requests that real applications will throw at it. The author asks for a public dataset of human-labelled image questions shaped like real app requests, and that is the right ask. Without that, every accuracy number here is a proxy for something unverified.
The second open question is more practical: has anyone run SmolVLM in MLX for scoring rather than generation? The author notes that the current latency on the Mac is acceptable but that a larger backbone would be more affordable if MLX supported prefix sharing for scoring. That is a concrete, actionable request to the community, and it points to the real bottleneck in on-device AI: not raw compute, but software that supports the clever inference tricks. We would tell a reader considering this tool to try it, but to treat its probabilities as well-calibrated guesses until a human-labelled benchmark exists. The design philosophy, answers you asked for, in the format you asked for, fast enough to feel instant, is exactly what we should expect from AI-native tools. The missing piece is trust in the labels, and that is a dataset problem, not a model problem. Watch whether the author finds that dataset, because if they do, this 500M model becomes a quiet workhorse for every app that needs a quick visual decision without the bloat of a chatbot.