Testing LVFace against ArcFace for real-world face recognition speed

I'm considering transitioning from my current face recognition stack, which utilizes SCRFD and ArcFace models, to LVFace, based on the ByteDance paper from ICCV 2025.

3 min readMachine Learning

**Our Take: LVFace Looks Promising, But the Real Question Is Whether the Trade-Off Is Worth It for You**

The user is asking the right question, and we think the answer is more nuanced than a simple yes or no. LVFace, ByteDance's Vision Transformer-based face recognition model from ICCV 2025, clearly delivers on accuracy, especially for masked faces, where ArcFace has a well-known blind spot. That first-place finish in the MFR-Ongoing challenge isn't meaningless. But the core tension here isn't about benchmark scores. It's about whether that accuracy gain justifies the compute cost in a real, long-running production environment.

Let's be direct about what LVFace solves. ArcFace treats a mask like a confusing part of the face. It tries to compute embeddings for covered areas and ends up with noise. LVFace, by design, appears to handle this differently, focusing on the peri-orbital region and weighting the embedding more intelligently. That's a genuine improvement for anyone running recognition on people wearing masks, sunglasses, or partial occlusions. If your gallery is full of faces that are rarely fully visible, LVFace is likely to give you better recall with fewer false positives. The user's concern about embedding drift at scale is also valid, but the ViT architecture tends to produce more stable embeddings than CNNs, so we'd expect that to be less of an issue than with ArcFace.

The real friction point is inference speed and VRAM. ViTs are heavier than ResNet-50, and that's not negotiable. The user is running local inference with high-concurrency batching, which means every millisecond and every megabyte of VRAM matters. If you're serving thousands of requests per second, the extra compute overhead of LVFace could push your hardware costs up meaningfully. The trade-off is clear: you get better discrimination on occluded faces and potentially better recall at scale, but you pay for it in latency and memory. There's no free lunch here.

Our opinion is that LVFace is worth the swap if your use case is dominated by partially occluded faces or if your current false positive rate on masked faces is hurting your pipeline. If most of your faces are fully visible and your ArcFace pipeline is already well-tuned, the accuracy gain may not be large enough to justify the compute hit. Test it on your own data, with your own batch sizes, before committing. The user's small-scale testing is a good start, but production behavior under load is the only benchmark that matters.

From Machine Learning

I’m looking at swapping my current face recognition stack for LVFace (the ByteDance paper from ICCV 2025) and wanted to see if anyone has real-world benchmarks yet.

Currently, I’m running a standard InsightFace-style pipeline: SCRFD (det_10g) feeding into the Buffalo_L (ArcFace) models. It’s reliable, and I've tuned it to run quickly and with predictable VRAM usage in a long-running environment, but LVFace uses a Vision Transformer (ViT) backbone instead of the usual ResNet/CNN setup, and it supposedly took 1st place in the MFR-Ongoing challenge.

Read the original at Machine Learning