There's a particular kind of joy in watching someone shrink a cultural artifact down to its mathematical bones, and this Bad Apple experiment delivers exactly that. A 3.2 MB neural network, 790,000 parameters, has learned to reconstruct a heavily subsampled version of the classic animation, frame by frame, from nothing but a 3D coordinate and a sine activation function. The source material was reduced to about a tenth of its original pixels, then distilled into weights. The result isn't a compression win in the traditional sense, and the author admits it openly: the subsampled video is 700 KB, while the network that reproduces it is larger. But that's not the point. The point is what it takes to make a neural network memorize a temporal sequence of images with enough fidelity to be recognizable, and the journey there is full of lessons that extend far beyond this one internet meme.
The technical path is genuinely instructive. The first attempt, a ReLU MLP with Fourier features, plateaued at a mean squared error around 0.12. The switch to SIREN's sine activations brought high-frequency detail for free, but introduced a new problem: the network could represent fine textures yet struggled to move information quickly, so fast motion came out blurry. Two fixes did the heavy lifting. Scaling the time coordinate by 4x relative to space gave the model more temporal capacity, and a motion-focused sampling strategy, pulling half of each batch from pixels that actually changed between frames, solved the starvation problem caused by Bad Apple's predominantly static black background. The result was a 9x drop in validation MSE, with high-motion frames 3.6x closer to ground truth and static frames nearly 15x closer. This is the kind of iterative, honest problem-solving that makes for genuinely useful reading, and it connects directly to broader questions about how we evaluate generative models. If you're building forecasting systems, the Beyond MSE: Refining Forecasts with Autoregressive Rollout and Uncertainty piece explores a similar tension between raw error metrics and perceptual quality, and this experiment is a vivid reminder that the loss function you choose shapes what the model learns to care about.
What stands out here is the deliberate, almost stubborn pragmatism. This isn't claimed to be a new frontier or a breakthrough. They're sharing a well-documented attempt with clear failure modes, and that's rarer and more valuable than another hype-driven demo. For anyone working with implicit neural representations, the decision to abandon per-frame finetuning after encountering catastrophic forgetting, and instead use a single shared network with a cosine-scheduled Adam and a final low-learning-rate polish pass, is the kind of detail that saves weeks of frustration. The takeaway is concrete: if you're training a network to memorize a signal with both static and dynamic regions, uniform sampling will starve the edges. You need to actively bias your batches toward change. That's not a niche trick; it's a principle that applies to any imbalanced temporal dataset, from sensor streams to video compression. And it's worth asking whether similar biases are missing from your own training pipelines.
The honest question this raises is about the future of neural video storage. Right now, the network is larger than the subsampled source, so there's no practical win in storage terms. But plans to try smaller models and full-resolution training are already underway, and that's where the real curiosity lies. If a model can be trained to reconstruct video from coordinates alone, what happens when you push parameter count down to the point where it genuinely competes with codecs? That's not a question for today, but it's a thread worth pulling. The next step to watch isn't whether the animation looks perfect; it's whether the architecture can generalize from memorization to interpolation, whether it can invent frames that were never in the training set. That would be the actual transformation. Until then, this is a well-crafted lesson in the value of understanding your data's structure, and a reminder that sometimes the best way to learn about a tool is to use it on something you love.
