What Palette Atlas has built is the kind of quiet engineering breakthrough that matters more than most flashy AI demos. By solving the problem of searching 48,695 public-domain paintings by color scheme, a task that should be computationally prohibitive, the project demonstrates how thoughtful approximations can unlock entirely new ways to explore cultural data. This is not about faster spreadsheets, but it shares the same DNA as the work we see in projects like Whitetree: smarter nearest-neighbor search for streaming sensor data and Explore cost-effective vector search with on-disk ANN indexes like DiskANN. All three are asking the same question: how do we make similarity search practical at scale?
The core insight here is elegantly simple. The natural distance between two paintings' color palettes is the Earth Mover's Distance (EMD), but running that calculation for every pair across nearly 50,000 images would take hours. The team behind Palette Atlas found a shortcut: project the colors onto eight directions on a Fibonacci hemisphere, sort them, and average into 16 quantiles. The resulting 128-dimensional vector can be compared using simple L1 distance, which acts as a lower bound on the true EMD. The results speak for themselves, 78% overlap with exact top-10 results, and the top-10 that differ are only 1.8% farther in true distance. Compare that to a soft color histogram approach, which achieves only 45% overlap and 13.5% error, or a naive mean-color method that is essentially useless at 3% overlap and 161% error. The sliced embedding is not perfect, but it is dramatically better than any obvious alternative.
What makes this project particularly impressive is the attention to engineering detail. The index uses a Hierarchical Navigable Small World (HNSW) graph written from scratch, achieving perfect recall@10 at query time with ef=32 in just 0.25 milliseconds, 56 times faster than brute force. The vectors have an intrinsic dimension of about 11, meaning the hierarchy actually helps rather than hinders, a nuance recently explored in the Munyampirwa et al. 2025 paper that Palette Atlas explicitly cites. The top layer stores eight k-means landmarks instead of randomly selected nodes, turning the entry point into a curated gallery of one typical painting per color family. This is not a cosmetic tweak; it means the first node you encounter is already representative of a coherent color cluster, which makes the search both faster and more interpretable.
The practical takeaway here is direct: if you are building any system that needs to match visual or perceptual data by similarity, the sliced-Wasserstein approach is worth studying. It is not a general replacement for learned embeddings, but it is a principled, mathematically grounded alternative that does not require training data or GPU inference. The code and write-up are open source, meaning anyone can adapt this technique for their own domain, whether that is searching product catalogs by color palette, matching fabric swatches in design tools, or even comparing sensor signatures in streaming data, as Whitetree explores. The question Palette Atlas leaves open is whether this approach generalizes beyond color to other perceptual dimensions like texture or shape. That is the next frontier, and the foundation is already laid.