PCA rotation turns ordinary embeddings into compressible ones

In exploring the effectiveness of PCA before truncation on non-Matryoshka embeddings, I tested a method that enhances compressibility.

3 min readMachine Learning

There's a quiet assumption buried in most embedding workflows: if a model wasn't trained for Matryoshka-style truncation, then compressing its output means accepting serious damage. The numbers shared here challenge that assumption directly. Fitting PCA once on a sample, rotating the vectors into that basis, and then truncating turns a 1024-dimensional BGE-M3 embedding into 128 dimensions while keeping 0.933 cosine similarity. Naive truncation at the same width collapses to 0.333. That's not a marginal improvement. It's the difference between a method that's broken and one that's genuinely usable.

The practical implication is straightforward: you don't need a new model to get compact embeddings. You need a better linear transform. PCA is not exotic. It's a standard tool that most practitioners already have in their stack, and it costs a single fitting pass on a representative sample. After that, truncation stops being an arbitrary slice and starts being a principled way to concentrate signal into the leading components. For teams running non-Matryoshka models in production, this could mean smaller vector indexes, faster retrieval, and lower memory overhead without retraining or swapping to a different embedding model. That's a meaningful unlock for a lot of real-world systems.

That said, the evaluation raises a point worth sitting with. Cosine similarity is forgiving. Recall@10 is not. At 27.7x compression with PCA plus 3-bit quantization, cosine still reads at 0.979, but recall drops to 76.4%. That's a solid result, but it's not the same as the cosine score. If your application depends on retrieving the right item, not just scoring well on a similarity metric, then the aggressive end of the spectrum will cost you. The piece rightly flags this. Too many compression discussions stop at "look how close the vectors are" without asking whether the ranking survives. For retrieval, ranking is the product.

The open questions here are the right ones. Is PCA the best linear baseline, or is there something stronger that doesn't require training a full model? Which metric should dominate decisions: reconstruction fidelity, recall, or something task-specific? And do these patterns hold across other non-Matryoshka architectures, or is BGE-M3 a favorable case? Those aren't rhetorical. They're the questions that determine whether this technique becomes a footnote or a standard practice. The method deserves serious engagement, but it also deserves scrutiny. Compression is only useful when it preserves what your application actually needs. The data here suggests PCA-first truncation is a strong candidate. It's not a free pass to push compression as far as possible and assume the recall will follow. Test it on your data, with your metric, at your operating point. That's the only way to know if it works for you.

From Machine Learning

Most embedding models are not Matryoshka-trained, so naive dimension truncation tends to destroy them.

I tested a simple alternative: fit PCA once on a sample of embeddings, rotate vectors into the PCA basis, and then truncate. The idea is that PCA concentrates signal into leading components, so truncation stops being arbitrary.

Read the original at Machine Learning