2 min readfrom Machine Learning

The qlora 2e-4 default is wrong under 10k samples and nobody talks about it [D]

Our take

Fine-tuning QLoRA models on smaller datasets—less than 10,000 samples—often leads to unexpected results. The pervasive default learning rate of 2e-4, widely promoted across tutorials and documentation, can actually trigger overfitting. Extensive experimentation reveals that a starting learning rate of 1e-4 or lower, combined with increased epochs, consistently yields significantly improved evaluation metrics. This adjustment, easily implemented, can save practitioners considerable time and frustration, as detailed in a recent discussion about ECCV expenses.

The frustration detailed in /u/Pretty-Ad774’s recent Reddit post regarding QLoRA fine-tuning resonates deeply with many in the AI community. The assertion that the default learning rate of 2e-4 is a “trap” for smaller datasets (under 10k samples) is a pointed critique of a widely accepted practice, especially considering the prevalence of tutorials and documentation that rigidly prescribe this value. The issue, as the author highlights, stems from the original Alpaca dataset's size (52k samples), a scale vastly different from the more modest datasets many practitioners are working with. This disconnect leads to overfitting, a frustrating cycle of decreasing training loss and stagnant or worsening evaluation loss, a scenario many of us have likely experienced—perhaps not quite as dramatically, but with similar wasted effort. This echoes anxieties shared within the research community about accessibility and cost, as seen in discussions about the prohibitive expenses of attending conferences like ECCV [Why is ECCV so insanely expensive for students presenting papers? [D]].

The core of the problem lies in the uncritical adoption of parameters derived from larger, different contexts. It's a common pitfall in rapidly evolving fields like AI – we latch onto what's presented as “best practice” without sufficiently questioning its applicability to our specific circumstances. The author’s account of spending weeks troubleshooting, only to find a solution in a seemingly minor adjustment to the learning rate (down to 1e-4), underscores the importance of empirical experimentation and a healthy dose of skepticism. The fact that Unsloth’s documentation labels 2e-4 as a “starting point” while numerous tutorials hardcode it is particularly galling, contributing to a cycle of misinformation and wasted resources. This situation highlights a broader challenge regarding documentation and knowledge sharing in the AI space, particularly as complex techniques like QLoRA become more democratized, as concerns about the reliability of AI agents for enterprise deployment also demonstrate [Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026]. The process of peer review and rigorous validation of methodologies often lags behind the initial hype and rapid release of tools and techniques.

The implications of this discovery extend beyond just QLoRA fine-tuning. It serves as a broader reminder of the need for data-driven parameter selection and a more nuanced understanding of the interplay between dataset size, model architecture, and training hyperparameters. While there's a temptation to seek universal solutions, the reality is that optimal configurations are often context-dependent. This reinforces the value of individual experimentation and critical evaluation of established guidelines. The author’s proposed rule of thumb – 2e-4 for datasets above 30k, 1e-4 or lower for datasets under 10k, and active tuning in between – offers a practical starting point, but should ultimately be validated through empirical testing. Similarly, the recent debates around TACL journal acceptance rates [TACL journal doubts [D]] highlight vulnerabilities in relying on established processes without rigorous scrutiny.

Ultimately, this situation prompts a crucial question: how can we foster a culture of more critical evaluation and experimentation within the AI community? Should documentation prioritize providing ranges and guidelines for parameter tuning rather than prescriptive defaults? Could platforms like Hugging Face incorporate more robust data size considerations into their example notebooks? As AI development becomes increasingly accessible, the onus shifts from blindly following instructions to developing a deeper understanding of the underlying principles and the ability to adapt them to specific use cases. The future of AI innovation might well depend on our ability to move beyond the "copy-paste and hope" approach and embrace a more iterative, data-informed methodology.

Every qlora tutorial on earth says start at 2e-4. Unsloth docs, hf examples, the paper itself. and for small datasets i now think that numbers is a trap.

Where does 2e-4 come from? alpaca. 52k samples. cool, except most of us are fine tuning on 5-10k samples we scraped and labeled ourselves, not 52k. at that size the model overfits inside epoch one and then youre just watching training loss go down all pretty while eval lost sits there doing nothing. or climbs.

I burned close to three weeks on this. recleaned the data set twice. rewrote the prompt template twice. spent on entire sunday hand relabeling rows while my flatmate watched football next to me (started with 8k rows, ended around 7200 after cutting garbage, i think, didnt log it properly). eval did not move. you you know what’s worse than a bad eval? seven identical bad evals in a row.

Then i changed one number. 2e-4 down to 1e-4, epochs 3 to 5. eval jumped more than everything else combined. i sat there refreshing wandb thinking it was a fluke. three more runs, same story.

And the annoying part, unsloth literally calls 2e-4 “a starting point” in their own docs. but every shared notebook has it hardcoded, zero comment. so people copy paste, get garbage, blame their data, blame their rank, lose a week. ask me how i know lol.

My rule now. above 30k 2e-4 is probably fine. under 10k, start at 1e-4 or lower and add epochs. in between, actually tune it, its one number, takes an afternoon.

If there’s real research defending flat 2e-4 on small data i want to read it. and if you all quietly figured this out in 2024 and never posted about it, im mad at every one of you individually.

submitted by /u/Pretty-Ad774
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article