There is a quiet assumption buried in every fine-tuning tutorial, one that treats a single learning rate like a universal constant rather than a single data point from one specific dataset. The 2e-4 default that shows up in Unsloth docs, Hugging Face examples, and the QLoRA paper itself traces back to Alpaca's 52,000 samples. That number worked there. But most of us are not training on 52,000 carefully curated examples. We are scraping together five or ten thousand rows, cleaning them on a Sunday afternoon, and then wondering why our evaluation loss flatlines while training loss keeps dropping. The author of this post spent three weeks in that exact loop, recleaning data, rewriting prompt templates, and hand-labeling rows while their flatmate watched football. The fix was one number: drop 2e-4 to 1e-4, add a couple of epochs, and suddenly the eval moved more than all previous runs combined.
This is not an exotic edge case. This is the reality for a large portion of people fine-tuning open models today, and it connects directly to a broader lesson we have been circling for a while: the quality of your input pipeline determines more than any hyperparameter ever will. In our piece on Clean Data Starts With Catching AI Slop Before It Skews Your Model, we saw how filtering out AI-generated text made a sentiment model less accurate, because the label noise was carrying real signal. The same principle applies here. If your eval is stuck, the first instinct is to blame the data, and sometimes that is fair. But this experience suggests we should also interrogate the settings we copied without thinking. The default learning rate is not a law of nature; it is a starting point for one dataset. Treating it as a constant means you are optimizing the wrong variable. And when you are working with small data, the interaction between learning rate and epoch count becomes the whole game. Too high a rate and you overfit inside epoch one. Too aggressive a schedule and you never give the model room to generalize.
The frustrating part is that the tools are not hiding this. Unsloth's own documentation calls 2e-4 a starting point, but shared notebooks hardcode it with zero comments, so people copy, paste, and then spend a week blaming their labels or their rank. A refreshingly practical rule of thumb: above 30,000 samples, 2e-4 is probably fine. Under 10,000, start at 1e-4 or lower and add epochs. In between, tune it yourself, it is one number, takes an afternoon. That is the kind of concrete, actionable guidance that actually empowers people. It also echoes what we explored in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where practical constraints often matter more than theoretical elegance. The same logic applies to fine-tuning: the best recipe is the one that matches your data size, not the one that worked for a famous benchmark.
The open question this raises is whether the research community will ever publish a systematic study of learning rate and epoch interactions across dataset sizes for QLoRA, or whether we are all expected to rediscover this through trial and error. The author would read that paper. We would too. But until it exists, treat every default as a suggestion, not a rule. Run a small grid, log your eval, and trust the curve over the copy-paste. That is the difference between following a tutorial and actually understanding your model.