Slimming face embeddings to fit PostgreSQL with 16-bit precision.

ArcFace embeddings, when quantized to 16-bit using pgvector HALFVEC, present an intriguing opportunity to optimize storage and I/O operations in PostgreSQL.

3 min readMachine Learning

The math here is refreshingly simple, and it works. Dropping ArcFace embeddings from 32-bit to 16-bit precision is not a hack; it is a pragmatic response to a real constraint in PostgreSQL. When a 512-dim vector plus its header crosses the 2040-byte TOAST threshold, you are not just paying a storage premium. You are forcing the database to take an extra read on every row that does not live in the main heap. Halving the precision to HALFVEC brings that payload under the line, which means the vector stays inline, gets read in the same page fetch, and effectively halves your I/O for the same logical query. That is not a marginal optimization. That is the difference between a system that feels responsive and one that stalls on disk lookups.

The more interesting question is whether you lose anything meaningful in accuracy. Trusting the training dynamics here is the right call. ArcFace and similar loss functions are explicitly designed to push same-identity embeddings close together and different-identity embeddings far apart, often with a large margin baked into the loss itself. That separation is not delicate. It is robust to the kind of noise you introduce by quantizing from 32-bit to 16-bit floats. The reported difference in a normalized 0-to-100 similarity score lands around the third decimal place, roughly 0.001. For any real-world face matching task, whether that is clustering a photo library or verifying an identity against a watchlist, that level of drift is imperceptible. You are not trading a meaningful drop in precision for a storage win. You are giving up noise that was never serving you.

This is also a reminder that precision is not a virtue on its own. The industry has a habit of treating 32-bit floats as the default because that is what the model was trained on, but the model was trained on a loss that creates wide margins, not on the last bit of floating-point accuracy. If the embedding space is structured to be separable, the quantization error is just another form of input noise, and the model already has enough slack to absorb it. The practical takeaway is that you should test this on your own data, but the prior should be strongly in favor of 16-bit working just fine. You are not the first person to try this, and you will not be the last, because it is the right call for most production workloads.

So the answer to the original question is yes, this sounds right, and yes, it is a standard approach. The real insight is that the TOAST threshold is not just a storage detail. It is a performance cliff that your vector type can either avoid or fall off. By going 16-bit, you are not just saving bytes. You are keeping your queries in the fast path, and that is worth more than a few decimal places of theoretical accuracy. Run the benchmark, confirm the similarity scores, and then ship it.

From Machine Learning

512-dim face embeddings as 32-bit floats are 2048 bytes, plus a 4-8 byte header, putting them just a hair over over PostgreSQL's TOAST threshold (2040 bytes), meaning by default postgresql always dumps them into a TOAST table instead of keeping them in line (result: double the I/O because it has to look up a data pointer and do another read).

Read the original at Machine Learning