CPU Inference
CPU Inference at Beyond Market Intelligence is a file of 2 stories. The newest of them: “Smaller model, sharper reasoning: 44M parameters trained from scratch on CPU” and “Rebuilding a Smarter Language Model from the CPU Up”. Three weeks ago, SHADOW-250M proved a 60 MB model could retrieve records from disk. After six months away, developer zemondza is rebuilding NORD, their spiking language model, with a sharper focus: CPU-first inference from the ground up. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work… The list below is every CPU Inference story on Beyond Market Intelligence, newest first.
Smaller model, sharper reasoning: 44M parameters trained from scratch on CPU
Three weeks ago, SHADOW-250M proved a 60 MB model could retrieve records from disk. Now, SHADOW-50M pushes the same idea further: 44M parameters, 19.8 MB, and roughly 1,900 tokens per second on a laptop CPU. It reasons over what it retrieves, runs arithmetic through a fixed circuit, and answers from memory without re-reading source text. It loses to larger models on standard benchmarks, and the author says so plainly.
Rebuilding a Smarter Language Model from the CPU Up
After six months away, developer zemondza is rebuilding NORD, their spiking language model, with a sharper focus: CPU-first inference from the ground up. The new version, NORD 5.5 Flash, drops the artificial spike-time dimension in favor of using the actual token sequence as the time axis. That's a cleaner design choice. The goal isn't to outrun Transformers, but to see if a simplified, truly causal architecture holds its own. We're curious to see the benchmarks when they arrive.