Scale preprocessing without the DevOps overhead

Managing long-running preprocessing jobs at scale can be a significant challenge, especially for small machine learning teams.

3 min readMachine Learning

The wall that u/krishnatamakuwala describes is not a technology problem. It is a design problem hiding inside a tooling question. When a small ML team spends more time wrestling with job orchestration than with preprocessing logic, the issue is not that Prefect or Temporal are bad tools. It is that those tools were built for teams with dedicated infrastructure engineers, and this team does not have one. That distinction matters because it shifts the conversation away from "which framework should we adopt" and toward "what should preprocessing look like when the team's expertise is the model, not the pipeline."

The practical reality for most small ML teams is that the biggest failure point is not the single machine crashing halfway through a 50GB job. It is the overhead of maintaining the distributed system that was supposed to fix that crash. Every hour spent configuring a worker pool or debugging a DAG is an hour not spent improving the model. The team has already seen this firsthand: they looked at Prefect, Temporal, and others, and correctly identified that each one demands a level of operational maturity they do not have. That is not a failing of the team. It is a mismatch between the tool's assumptions and the team's actual constraints. The honest answer to their question about whether distribution is worth the setup overhead is almost always no, unless the preprocessing job is so large that a single machine cannot hold it at all.

What is actually working for teams in this position is not a better orchestrator. It is a preprocessing approach that treats the data as a series of independent chunks rather than a single monolithic pipeline. If a 50GB dataset can be split into 50 one-gigabyte shards, each shard can fail independently without destroying the entire job. That pattern does not require a distributed system. It requires a simple retry loop and a checkpointing strategy that the team can write in an afternoon. The failures still happen, but they stop being catastrophic. The team can resume from the last successful shard instead of starting over. That is the concrete improvement they should pursue: not more infrastructure, but less coupling in the data itself.

The takeaway here is straightforward. Stop searching for a tool that will abstract away the complexity of distributed preprocessing. That tool does not exist for a team without DevOps support. Instead, redesign the preprocessing to be resilient at the data level. Split the work into independent units. Build a simple retry mechanism. Accept that failures will happen, but make them cheap. That is how small teams scale preprocessing without scaling their operational burden.

From Machine Learning

We're a small ML team for a project and we keep running into the same wall: large preprocessing jobs (think 50–100GB datasets) running on a single machine take hours, and when something fails halfway through, it's painful.

We've looked at Prefect, Temporal, and a few others — but they all feel like they require a full-time DevOps person to set up and maintain properly. And most of our team is focused on the models, not the infrastructure.

Read the original at Machine Learning