If you have been watching the slow creep of memory bottlenecks in inference, this 12-bit lossless BF16 format from a lone Australian researcher is the kind of work that deserves more than a passing glance. The claim is direct: store weights in 12 bits instead of 16, lose nothing in precision, and decode with a single integer ADD for 99.97% of values. That is not a marketing promise. It is a measurable property of the format, and it holds across models as varied as Llama 3.1 405B and Mixtral 8x7B, with escape rates well under a tenth of a percent.
What makes this practical is not the compression ratio alone. A 1.33x reduction in memory footprint is useful, but the real win is that the format is designed for direct use during inference. The sign and mantissa live in one byte, the group code in another, and the decode fuses into the matrix multiplication kernel. That means no separate decompression stage, no lookup tables, no bitstream parsing. The numbers reflect that design: single-user throughput on a 5070 Ti reaches 64.7 tok/s for Llama 2 7B, and multi-user scaling is even more striking, with 2931 total tok/s versus 1086 in vLLM. Those are not incremental gains. They are the kind of results that make you reconsider what is possible on consumer hardware.
There is a temptation to dismiss this as a niche trick or a research curiosity, especially since it has only been tested on BF16 safetensors so far. But the stability across model families, from diffusion UNets to video generation, suggests the underlying idea is not fragile. The escape rate, the small fraction of weights that need extra handling, stays low enough that the decode path remains simple. And the fact that it works on both NVIDIA and AMD means the approach is not locked into a single vendor's toolkit. That is the kind of portability that matters when you are building for real deployments, not just benchmarks.
The honest question is whether this gains traction beyond the prototype stage. The new format invites criticism and edge cases, and that is the right posture for work that challenges assumptions about what lossless compression can do. But the evidence so far is compelling: byte-aligned storage, zero read amplification, and a decode cost that is effectively free when fused into GEMM. If this holds up under broader testing, it is not just a faster way to run existing models. It is a reason to rethink how we store and move weights in the first place. For anyone who has ever watched a model stall on memory bandwidth, that is a future worth exploring.
