SGD

3 stories filed under SGD on Beyond Market Intelligence. The newest of them: “Peek inside a neural network as it learns, layer by layer”, “When a scientific calculator becomes a platform for AI discovery”, and “Train smarter by isolating data reuse bias in gradient descent.”. Open the training view and you'll watch a small neural network learn to read digits, layer by layer. A 67% accuracy on handwritten digits sounds unimpressive, until you remember the entire model fits in your head. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work… The list below is every SGD story on Beyond Market Intelligence, newest first.

Peek inside a neural network as it learns, layer by layer
Machine Learning

Peek inside a neural network as it learns, layer by layer

Open the training view and you'll watch a small neural network learn to read digits, layer by layer. Built with plain NumPy and manual backprop, this tool exposes the messy, instructive details most courses skim over: gradient norms, inactive neurons, and weight shifts from initialization. It reaches about 98.5% on MNIST, but the real value is the lab, where you can ablate neurons or tweak temperature and see accuracy move instantly. For teachers and self-learners, it turns abstract theory into something you can touch.

Machine Learning

When a scientific calculator becomes a platform for AI discovery

A 67% accuracy on handwritten digits sounds unimpressive, until you remember the entire model fits in your head. This builder trained a binary classifier on the Casio FX-82CE X, hand-punching weights from just six images. The architecture is brutally simple: a 3x3 grid feeding one neuron. That it works at all is a quiet win for minimalism. The real lesson? Even a clunky, manual process reveals how much signal lives in raw data.

Train smarter by isolating data reuse bias in gradient descent.
Machine Learning

Train smarter by isolating data reuse bias in gradient descent.

Training a neural network until the training error vanishes while the test error stagnates is a familiar frustration. *Decoupled Descent* frames this as data reuse bias, isolating it with full-batch gradient descent on Gaussian mixture models. By applying approximate message passing corrections, the method guarantees that training and test errors asymptotically align at every iterate. It's a theory paper, so large-scale models remain a stretch, but the certificate it offers is a thoughtful step toward principled stopping.