font embedding

Uncovering Hidden Structure: How Font Embeddings Bloom Into Visual Patterns

When Dylan Berndt fed thousands of Google Fonts through a custom neural network and compressed the resulting embeddings into color and position, something unexpected bloomed: a flower.

4 min readMachine Learning
Uncovering Hidden Structure: How Font Embeddings Bloom Into Visual Patterns
Embedding Every Font with Neural Networks makes some Nice Structures (including a flower) [P]

Gubi reissued the Lotus lamp. That's the kind of detail that surfaces when you stare at a visual map long enough, which is exactly what Dylan Berndt has built with Font Search. His project transforms thousands of fonts into a landscape of colored dots, each one an embedding generated by a custom neural network. When he ran the Google Fonts corpus through this process, the resulting structure bloomed into something organic: a flower, with cursive fonts clustering naturally in the stamen and serif families forming the petals. This is not a gimmick. It reveals something fundamental about how visual structure can emerge from data without being forced. We think this is exactly the kind of exploration that matters most in an AI-native world, because it points to a future where tools find meaning for us rather than waiting for us to specify every parameter. You can see echoes of this approach in how a model trained on a billion chess positions can reshape data analysis, which we covered in Explore a Billion Chess Positions to Transform How You Analyze Data. Similarly, the way language models learn to copy context with hash tables, detailed in How language models learn to copy context with hash tables, shows a parallel logic: the network discovers hidden patterns on its own, and the human's job is to notice what it found.

The practical consequence for anyone working with spreadsheets or data tools is immediate. Traditional font search relies on tags, names, or manual categories, designers spend hours filtering by "sans serif" or "display" and still miss the perfect match. Berndt's embedding approach does what spreadsheets cannot: it understands visual similarity as a continuous space, not a checkbox. A font might be halfway between slab serif and handwriting, and the map places it precisely where it belongs. That is the difference between a lookup table and a living structure. It is the same shift we highlighted in Explore how to filter 5000 IDs outside your Power BI model, where the challenge is not having the data but having the right way to navigate it. Berndt's tSNE projection, which he found outperformed PCA and UMAP, is a deliberate choice that preserves the flower pattern rather than flattening it into something generic. That matters. It means the method respects the data's own shape, and the result is a map you can actually explore with intuition instead of queries.

What stays with us is a specific detail from Berndt's own commentary: he says the neural networks need a post-training step to adapt them for the actual task of searching, but the pre-trained models produced the most interesting structure. That is worth sitting with. The thing that works best was never optimized for the final task, it simply learned to see fonts, and what emerged was an unexpected bloom. For anyone building AI tools today, that is a concrete lesson: the most valuable output might come from the part of the pipeline you treat as a means rather than an end. The question Berndt leaves us with is not whether his font map is useful for searching, it clearly is, but what other hidden structures are waiting inside the datasets we already have, quietly arranging themselves into flowers we have not bothered to look at.

From Machine Learning

I've been working on a font searching tool for about a year now, and my investigations have centered around pre-training neural networks to produce embeddings of each font. I usually then need to post-train the networks to adapt them to the task of font searching. However, the most interesting thing I created in the course of the project came from the pre-trained models.

Read the original at Machine Learning