Vision-language models are not built from scratch at all. They are text-only language models that have been taught to process images as if they were just another language. That distinction matters, because it changes how we should think about what these models can actually do.
A training process treats visual information as a translation problem. The model already understands syntax, context, and meaning from its text training. The fine-tuning simply adds a visual encoder that maps image patches into the same token space the model already uses for words. It is elegant engineering, but it is also a practical limitation. The model does not see the way we see. It sees by converting pixels into tokens that happen to fit its existing language architecture.
For anyone working with data, this has immediate implications. If you are evaluating a vision-language model for a spreadsheet task, you should not assume it understands spatial relationships or visual layouts the way a human does. It understands them the way a language model understands a paragraph about a table. That can work well for structured data, but it will fail in ways that surprise you if you treat the model as a visual thinker. The vision is borrowed, not native.
We think this is the right framing for the industry to adopt. Too often, capabilities are described as if models have developed new senses. They have not. They have learned to translate one input format into another format they already know. That is powerful, but it is also bounded. The practical takeaway for spreadsheet users is straightforward: test your vision-language model on the actual documents you need it to read, not on idealized examples. The model will perform exactly as well as its training data allows, and no better. That is not a criticism. It is a reminder that tools work best when you understand what they are actually doing.
