The AI industry is chasing scale, but scale alone won't solve the fundamental inefficiency baked into every modern neural network. The post from /u/Sevdat points toward a genuinely different path, one that trades brute-force computation for architectural intelligence, and it deserves serious attention.
Here is the core insight: backpropagation is a bottleneck. It requires storing intermediate values, propagating errors backward through every layer, and repeating that process millions of times. Predictive coding offers an alternative where each neuron decides locally whether to activate, based on its own prediction error. Combine that with a 1-bit LLM architecture, where each weight is either +1 or -1, and you eliminate most of the memory and computation overhead. The result is a system that runs on calculated chance rather than deterministic gradient descent. For users, this means models that could run on far less power and memory, potentially on devices you already own, without the cloud dependency that currently limits privacy and speed.
The second move is equally practical: instead of expecting a single forward pass to produce the right answer, let the model re-prompt itself iteratively. Store context in RAM and let the model pull relevant information to adjust its weights dynamically for each query. This avoids catastrophic forgetting because the model isn't overwriting its core training, it's adapting temporarily per task. The efficiency gains from the 1-bit architecture make this loop viable where current hardware would choke on the latency.
The hardware question is the hard part. Modern chips are built for deterministic, sequential logic. A stochastic system needs hardware that treats randomness as a resource, not a bug. /u/Sevdat points to heat as a noise source and to Extropic's TSU as a rare example of someone attempting this. The physics of the metal itself could decide whether a neuron fires. That is not science fiction; it is engineering that nobody is funding at scale. Without it, the current AI bubble risks stagnation because scaling laws are already showing diminishing returns. The practical takeaway is clear: the next leap in efficiency will not come from bigger clusters or more data. It will come from rethinking the fundamental unit of computation, and building the hardware to match.