Discover AI-native compression that makes 7B models run smarter on Mac.

Introducing INT3 compression with fused metal kernels, a breakthrough for model optimization in AI.

3 min readMachine Learning

If you've been waiting for a reason to take local AI seriously, this is it. A solo researcher has shipped INT3 compression for a 7B model with a loss of just +0.14 nats, alongside a 2-bit KV cache designed for long-horizon tasks. That's not a demo. That's a working implementation, available now for Mac users via Homebrew, with Qwen 7B already in preview. The practical takeaway is straightforward: your M-series Mac can now run a capable model with a fraction of the memory footprint, and the trade-off in quality is small enough to matter only in the most sensitive contexts.

What stands out is not just the compression numbers, but the intent behind them. This isn't a research paper tucked away for peer review. It's a functional tool with custom fused Metal kernels, built specifically for Apple silicon. For anyone who has wrestled with running local models, that's meaningful. It means faster inference, lower memory pressure, and the ability to hold longer conversations or process longer documents without watching the context window shrink. The 2-bit KV cache is the quiet hero here. Most compression efforts focus on weights, but attention states grow quickly, and cutting those to two bits changes what's possible on a laptop.

We should be clear about what this does not claim. It doesn't promise a universal solution. The researcher notes there's more room to pack efficiently, and GPU support via Triton kernels is still in progress. That's honest. It also signals where the work is heading. The path from a 7B model to something larger, even up to 100B parameters, is the stated ambition. That's not hype. It's a roadmap, and one that invites feedback. The request for input on which models to compress next is a practical way to build something the community actually needs, rather than another benchmark chase.

The real point here is that the gap between frontier-scale models and what runs locally is narrowing faster than most people assume. You don't need a cluster to experiment with strong language models. You need better compression, tighter kernels, and a willingness to ship. This project has all three. If you're on a Mac and curious about what a 7B model can do without the cloud, the install command is one line away. That's the kind of frictionless start worth exploring.

From Machine Learning

Hey guys, I am a researcher and solo founder. I compress models with INT3 at +0.14 nats and built a 2-bit KV cache for long-horizon tasks. I shipped both (INT3 model + INT2 KV) with custom fused Metal kernels for Mac (M-series). Currently Qwen 7B is available in preview.

I am optimizing kernels further and working on Triton kernels for GPU support. There is still more room to pack more efficiently, I will share more models soon. I will appreciate any feedback or any model you want me to compress within 100B parameters.

Read the original at Machine Learning