boilerplate

Stop rewriting the same spreadsheet logic. Let AI handle the pattern.

The developer's instinct to cut setup time from three days to one is smart, but the real insight here is the question they're circling: why are we still writing code at all?

4 min readMachine Learning

Every time we start a new model, we rebuild the same scaffolding. The data validation checks, the feature transformation logic, the config parsing. It is nearly 80 percent identical to the last project, yet we treat it as a fresh act of creation. The developer behind this reflection tried cookiecutter templates first. That worked until the template drifted from reality, because nobody wants to maintain a repo that exists only to be copied. A shared library was better, but the glue code still introduced bugs. Now they are using generative code to produce the boilerplate, and it cuts setup from three days to under one. That is real progress. But it raises a question that deserves more than a passing thought: if the code writes itself, should we be writing it at all?

This is where the conversation gets interesting. The question is not how to write better code. They are asking whether the code is even the point. A config-driven approach feels cleaner, more declarative, closer to the data problem itself. But they admit the risk: the moment you need something non-standard, the config becomes a prison. The opinionated framework that saves you time in the short term becomes the thing you fight in month four. This is not a new tension, but generative tooling has sharpened it. When the boilerplate writes itself, the real work shifts to the parts that resist automation. That is where judgment lives. That is where the model's behavior diverges from what a template can capture. And that is why we need to be careful about optimizing for the first day of a project when the real cost is often in the maintenance that follows.

There is a parallel here to what we have seen with data quality and model drift. In Clean Data Starts With Catching AI Slop Before It Skews Your Model, the issue is that automated filtering introduced its own bias, and the model got worse because the data pipeline silently degraded. The same principle applies here. If you automate the scaffolding without understanding what it is doing, you inherit the errors at scale. Similarly, Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges shows that real-world systems are messy, constrained by hardware and context in ways that synthetic benchmarks ignore. The lesson is consistent: the closer you get to production, the more the edge cases matter. And the edge cases are exactly what generated code tends to miss.

So what is the middle ground? The middle ground is not a better template or a smarter library. It is a clearer separation between the parts that are genuinely repetitive and the parts that require reasoning. Use generative tooling for the boilerplate, but treat it as a draft, not a deliverable. Keep the config-driven approach for what it does well, but design for escape hatches from the start. And above all, recognize that the goal is not to eliminate code. It is to eliminate the time you spend re-typing the same logic and free that time for the problems that are specific to your data, your model, and your users. Skepticism of a silver bullet is warranted. But they are also right that three days of setup is too long. The question they should be asking is not whether to write code, but which code is worth writing by hand. That is the discipline that will age well.

From Machine Learning

I have been reflecting on this while working on a project recently. Every time we start a new model, we rewrite roughly same scaffolding, data validation checks, feature transformation logic ; all of this is nealy 80 percent identical to last project

I tired templating with cookiecutter style project generators. Initially it was okay, but it drifted from reality since noone wants to maintain a template repo. So tired a shared library approach, it helped and was much better. But weiting glue code to wite everything is still bug prone

Read the original at Machine Learning