Bad Apple

How a tiny AI system learned to generate Bad Apple without a clock

Teaching a network to *remember* a video is one thing.

4 min readMachine Learning
How a tiny AI system learned to generate Bad Apple without a clock
Generating Bad Apple autonomously from a single initial state using a tiny recurrent dynamical system (417k params) [P]

There is a particular kind of magic in watching a model spin an entire video out of a single, silent vector. The Bad Apple experiment making the rounds this week is a sharp reminder that we often overcomplicate what it means to be creative with machines. The author took a classic benchmark, stripped away the crutch of a timestamp, and asked a tiny recurrent system to carry the whole narrative on its back. That is not just a neat hack; it is a different philosophy about where memory lives in a model. We have spent so long feeding networks explicit coordinates and conditioning signals, and here is someone showing that a 417,000-parameter system can learn to dream the sequence forward from zero. For anyone who has been following the slow march of generative models, this feels less like a novelty and more like a quiet argument about the power of compression. The practical choices in the training pipeline are where the real insight hides. The learned latent teacher tables and the rollout horizon curriculum are not just training tricks; they are a manual for how to coax long-term stability out of a system that has no business remembering six thousand steps. The lowest training loss did not translate to the best autonomous rollout; the real battle was conditioning the dynamics through noise injection and acceleration penalties rather than scaling up parameters. That is a valuable lesson for anyone who thinks bigger models are the only answer. We would point a curious reader to our piece on [Sharing my ML learning repo, NumPy to Transformers, 5 months, daily commits, all notebooks public. [D]](/post/sharing-my-ml-learning-repo-numpy-to-transformers-5-months-d-cmu91wx2d04i95ngm1211p2j9) to see how foundational knowledge compounds, and to Variational Autoencoders (VAEs) Explained: From Theory to ELBO and the Reparameterization Trick to contrast how latent spaces are usually built with explicit probabilistic goals. This work sits somewhere between those worlds, using a deterministic latent state but treating the training objective as a dynamical systems problem rather than a reconstruction task. What stands out to us is the refusal to rely on modern conveniences. No attention, no skip connections, no normalization, no truncated BPTT. The author stripped the architecture down to a four-gate LSTM-style recurrence and a simple decoder, then spent the effort on making the training signal do the heavy lifting. That is a contrarian choice, and it works. The system runs at over 200 FPS on a 4080 with a peak VRAM footprint smaller than a single high-res image. For practitioners, this is a direct challenge to the assumption that you need a transformer with billions of parameters to generate coherent long-form content. The trade-off is real, though, and the author admits the decoder is the weak point now. We would tell a reader who asks us about this that the takeaway is not the video itself but the recipe for stability. The second-difference acceleration regularization, the perturbation noise, the momentum scaling during curriculum changes, those are the parts worth stealing for your own projects. If you are struggling with recurrent models that diverge after a few hundred steps, this is a concrete set of levers to pull, and the code is public. The open question here is whether this approach scales beyond a well-known silhouette animation. The run was not perfectly scientific, with changes made mid-training, and the decoder needs serious work. But the core finding, that a small recurrent system can autonomously generate over six thousand frames from a single initial condition, should make us pause. It suggests that a lot of what we think of as temporal complexity can be compressed into a surprisingly small dynamical core if you train it with the right pressure. The next step we are watching for is whether this technique can generalize to more varied and higher-resolution content, because if it can, the era of feeding a model a prompt and letting it unroll a full narrative from a single latent seed might be closer than the parameter arms race suggests.

From Machine Learning

A few weeks ago, I saw this post where the author trained a SIREN MLP to implicitly memorize Bad Apple as a coordinate function: (t, y, x) to pixel.

That got me curious about a slightly different formulation: instead of handing the network a timestamp t, could a small recurrent dynamical system (RNN-ish) learn the continuous temporal flow in latent space and generate the entire ~6,500-frame full resolution video autonomously from a single initial condition (h_0, c_0)?

Read the original at Machine Learning