Scaling a Spiking Neural Network to 1.08B Parameters: and What It Reveals

In a bold exploration of Spiking Neural Networks (SNNs), I scaled a model to 1.088 billion parameters from scratch, driven by a desire to push boundaries in language modeling. Despite facing budget constraints and…

4 min readMachine Learning

There's a moment in every emerging field when someone tries the thing the papers say can't be done, and they do it with limited resources and a stubborn refusal to accept the conventional wisdom. That's exactly what this 18-year-old indie developer has done with a 1.08-billion-parameter spiking neural network trained purely in the spike domain. The established playbook said you need ANN-to-SNN conversion or heavy distillation to get past vanishing gradients at that scale. Instead, they pushed through 27,000 steps, hit a loss of 4.4, and ran out of money before they ran out of ideas. That alone is worth paying attention to, not because the model is fluent or production-ready, but because it challenges a boundary we were told was fixed.

What stands out most isn't the loss number. It's the structural behaviors that emerged when the architecture crossed the 600-million-parameter threshold. The model maintained roughly 93% sparsity, meaning only about 7% of neurons fired per token, which is a massive computational advantage during inference. More intriguing is the spontaneous shift in routing: as the model scaled past 600M parameters, it moved 39% of its activation routing into the persistent memory module. It essentially learned that memory matters more at scale, without being explicitly programmed to do so. And then there's the cross-lingual emergence, where around step 25,000 it started generating structurally correct Russian text despite no targeted weighting in the dataset mix. These aren't just curiosities. They're signals that the model is discovering general principles about efficiency and memory allocation on its own, which is exactly the kind of behavior that makes you wonder what else we've been over-constraining in our current architectures.

For anyone working in this space, the practical takeaway is twofold. First, pure spike-domain training at scale is harder, but not impossible, and the constraints might actually be leading to more interesting solutions than the conversion pipelines we've settled for. Second, the sparsity and memory routing findings have direct implications for hardware mapping. The developer is asking about neuromorphic platforms like Loihi, and it's a fair question. A model that maintains 93% sparsity and routes a significant portion of its computation through a dedicated memory module aligns well with architectures designed for event-driven, low-power inference. But the honest answer is that we don't yet know how well it maps, because the community hasn't had a model like this to test on. That's part of why this work matters: it gives us something concrete to probe, critique, and build upon.

The limitations are real, and the developer doesn't hide them. The text is janky, the loss is high, and the training was cut short by financial reality. But those are constraints of time and compute, not proof that the approach is a dead end. The milestone here is that a billion-parameter SNN converged from random initialization in the spike domain at all. That's not a finished product. It's a proof point that the ceiling we assumed for spike-based learning might be lower than we thought, and a reminder that sometimes the most useful experiments are the ones that run out of funding before they run out of insight. We'd like to see the code, we'd like to see the training curves, and we'd especially like to see what happens when someone with more compute picks up where this developer had to stop. That's the next step, and it's a concrete one.

From Machine Learning

Hey everyone. I’m an 18yo indie dev, and I’ve been experimenting with Spiking Neural Networks (SNNs) for language modeling. A lot of papers (like SpikeBERT) mention that training 1B+ SNNs directly from random initialization fails due to vanishing gradients, so people usually do ANN-to-SNN conversion or distillation. I wanted to see if I could force it to converge purely in the spike domain. I had to stop at 27k steps because my wallet is literally empty lol, but the loss converged to 4.4.

Here are the most interesting things that happened:

Read the original at Machine Learning