From scratch to scale: how MLEs build models today

Navigating the world of machine learning can be daunting, especially when transitioning between roles and preparing for interviews.

3 min readData Science

**Our Take: The real work of machine learning isn't model building, and that's fine.**

The user who posted this is not an anomaly. They are describing the actual career of a practicing MLE in 2025: eight years of experience spanning feature engineering, off-the-shelf model integration, data prep, hyperparameter tuning, and now building backend infrastructure for LLMs. They are being asked in interviews to hand-code transformer layers from scratch or explain early fusion versus late fusion on a whiteboard. And they feel like they are failing. They are not. The gap they perceive is not a skill gap, it is a mismatch between what interview processes test and what the job actually demands.

Look at the work they describe. They built MCP servers. They managed context windows. They created safety guardrails and engineered prompts. They pulled pre-trained models from HuggingFace and GitHub, adapted them to image tasks, and made them production-ready. This is the work that matters. The vast majority of ML engineering today is not about implementing attention mechanisms from scratch. It is about knowing which model to pick, how to connect it to your data pipeline, how to evaluate its outputs in a real business context, and how to keep it running reliably at scale. The models themselves are increasingly commodities. The engineering around them is where value is created.

We understand the anxiety. The interview process has not caught up with the reality of the field. There is a persistent cultural hangover from the era when every ML engineer was expected to be a research scientist who could derive backpropagation on a whiteboard. That expectation ignores how most teams actually build. When was the last time a production model needed a custom fusion layer? When was the last time you couldn't find a pre-trained checkpoint that solved 90% of your problem? The user's instinct, to search UnSloth or HuggingFace first, is the right one. It is efficient. It is pragmatic. It is what senior engineers do.

So here is the concrete point: If you are preparing for interviews, do not spend your limited time trying to memorize the internals of a transformer. Instead, spend it articulating how you chose a model, how you validated it, how you deployed it, and how you handled the inevitable failures in production. That story is more honest, more valuable, and harder than any hand-coded attention layer. The user's experience is not a weakness. It is the actual shape of the field.

From Data Science

I have about 8 years of experience mostly in the NLP space although i've done a little bit of vision modeling work. I was recently let go so I'm in the midst of interview prep hell. As i'm moving further along in the journey, i'm feeling i have some gaps modeling wise but I'm just trying to see how others are doing their work.

Read the original at Data Science