Making AI research accessible with an 18x cheaper, open-source pipeline

Unlock the potential of Karpathy's autoresearch for just $0.44 by leveraging an innovative open-source pipeline on SageMaker Spot instances. This solution enables 25 autonomous ML experiments to run in parallel,…

3 min readMachine Learning

This project is the most practical demonstration of accessible AI research we have seen in months. The author took Karpathy's powerful but resource-hungry autoresearch pipeline and stripped away the assumption that you need an H100 sitting idle for eight hours. By running 25 experiments for $0.44 instead of the usual $24, they have proven that the barrier to entry for autonomous machine learning is not technical skill, but a willingness to work with transient cloud infrastructure. That is a meaningful shift for anyone who has felt priced out of modern research.

For most practitioners, the real value here is not the raw cost savings, but the architecture choices that made them possible. The HUGI pattern, Hurry Up and Get Idle, is a clever solution to a common problem: GPUs that sit waiting for the next experiment are burning money. By spinning up a SageMaker Spot instance only for the duration of a five-minute experiment and terminating it immediately, the pipeline eliminates idle cost entirely. The author also documented a critical lesson about Spot instance pricing: larger instances can be cheaper than smaller ones because of lower demand, and regional availability varies wildly. These are the kind of battle-tested details that turn a clever idea into a reliable tool.

The project also reveals a surprising truth about cheap hardware. Experiments on an L40S produced results that transferred to an H100 for production training. Architecture rankings and optimizer choices held up across the price gap. Only absolute learning rates needed re-tuning. This suggests that for many exploratory tasks, an L40S at $0.04 per experiment is a perfectly valid research platform. The constraint is not the GPU, but the five-minute experiment budget, which ruled out deeper models and larger batches. That is a useful constraint to know about upfront.

What we appreciate most is the transparency. The author documented every failure, from Flash Attention 3 crashing on Ada Lovelace GPUs to the batch size trap where increasing device batch size made results worse. The eight-chapter vibe coding tutorial includes the actual prompts used and the exact commands to reproduce each step. Total cost to follow along: roughly seventy cents. That is how you empower people to explore. The pipeline is open-source, the tutorial is free, and the only remaining question is whether you will take the two hours to run it yourself.

From Machine Learning

TL;DR: I built an open-source pipeline that runs Karpathy's autoresearch on SageMaker Spot instances — 25 autonomous ML experiments for $0.44 total (vs ~$24 on an H100). 4x parallel execution, 2.3x faster, 18x cheaper. Includes an 8-chapter vibe coding tutorial. GitHub

Karpathy's autoresearch is brilliant — an AI agent modifies training code, runs 5-minute experiments, keeps improvements, and repeats overnight. But it assumes you have an H100 sitting around for 8 hours. Most of us don't.

Read the original at Machine Learning