GPT-2 Small

Explore how AI sees Trump before context enters the picture

A single word, "Trump," behaves like two different neighbors depending on how you slice its embedding.

3 min readMachine Learning
Explore how AI sees Trump before context enters the picture
GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P]

There is something quietly profound hiding in a simple visualization of GPT-2 Small's token embeddings. The post examines the nearest neighbours of the token "Trump" in two ways: one where each coordinate is thresholded into discrete bins, and one where the original continuous values are kept. The difference is stark. Discretized, you get generic political company: Mitt, Hillary, Pelosi, Blair. Continuous, you get a tighter, more personal cluster: family members, staff, rivals, and presidents like Obama, Clinton, Bush, and Eisenhower. No prompts, no text generation, just the static geometry of learned embeddings. And yet, the model is already telling us something about how it organises the world before it ever reads a sentence.

This is not a niche curiosity. It is a window into the difference between approximation and nuance, a theme that runs straight through the practical work of building and cleaning models. When you discretize, you force the model into a simpler, more brittle representation. The continuous space preserves the messy, overlapping semantics that make language human. This is the same tension behind the challenge of Clean Data Starts With Catching AI Slop Before It Skews Your Model: the choices you make about representation and filtering are not neutral. They shape what the model sees, and what it can learn. In that piece, over-zealous filtering of flagged reviews made a sentiment model worse because the genuine signals were lost. Here, the thresholding does something similar: it strips away the specificity that makes the embedding useful.

What we find instructive is that the continuous neighbours are not random or surprising. They are coherent. Family, rivals, predecessors. That is not a bug; it is the model's learned relational structure doing its job. The discretized version, by contrast, feels like a caricature of politics, flattening a person into a category. For anyone working with embeddings, the takeaway is direct: the representation you choose is the lens through which your model interprets every query, every document, every prediction. This echoes the practical lessons from Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where deployment constraints often force simplifications that have real consequences for accuracy. The same principle applies here: compression and discretization are not free lunches. They trade fidelity for speed or interpretability, and you should know exactly what you are giving up.

If someone asked us what to make of this, we would say: stop treating embeddings as opaque blobs and start interrogating them. The fact that a small model like GPT-2 Small encodes this kind of relational nuance in its static table is a reminder that the architecture is doing more than pattern matching on surface text. It is building a structured, almost intuitive sense of how entities relate. The open question for practitioners is whether your pipeline respects that structure or destroys it. When you quantize, truncate, or aggressively filter, you are not just saving compute. You are making a semantic decision. The concrete detail to watch is how often your evaluation metrics reflect the continuous space versus the discretized one. Because if you are only measuring performance on the latter, you might be optimizing for the wrong reality.

From Machine Learning

This visualization looks at the token “Trump” in GPT-2 Small’s static embedding table, before attention or context is applied.

The top plot is a t-SNE projection of 32,070 alphabetic tokens with at least two characters.

Read the original at Machine Learning