AI workflows

Build AI workflows that survive production without slowing your experiments

Every AI workflow faces a quiet tension.

3 min readInfoQ
Build AI workflows that survive production without slowing your experiments

There is a quiet tension at the heart of every AI workflow, and Mateus Moury has put his finger on it. The same machinery that makes a pipeline durable enough for production, with every step persisted and distributed across nodes, is precisely what bogs down the fast, throwaway loop you need to judge an LLM's output. This is not a minor inconvenience. It is a structural trade-off that every team building AI features will eventually hit, and Moury's pattern offers a way to stop treating that trade-off as a law of nature. The insight here is that you do not need one workflow. You need two, and they should never share the same runtime.

This is where the practical wisdom lands. If you have been building AI systems for any length of time, you have felt the pain of a 40-minute run just to check whether a prompt tweak improved a response. The heavy lifting of crash recovery and state persistence is non-negotiable for production, but it is dead weight when you are iterating on a single example. This pattern separates these concerns cleanly, and that separation is what we would tell any reader to steal immediately. It reminds us of the broader lesson that the right architecture is often about knowing what *not* to share. Related thinking on distributed systems and AI training reinforces this point, as seen in Unlock LLM Training: A Practical Guide to Distributed Algorithms, where the fundamentals of coordination matter more than any single framework. And when we consider how quickly we trust an AI's output, the need for fast evaluation loops becomes even more critical, a theme echoed in Verify Your AI's Understanding: A Simple Check for Tax Season.

The honest take is that most teams will read this and think they need to build a custom orchestrator. They do not. What they need is the discipline to recognize that durability and speed are not competing priorities to balance, but distinct concerns that demand distinct tools. The moment you accept that a single workflow runtime cannot serve both masters, you free yourself to use a lightweight evaluator for development and a heavy-duty runner for production. That is not a compromise. That is good engineering. We would caution against the temptation to abstract this into a platform feature before you have felt the pain on a real project. Start small. Let the pattern emerge from the work.

The specific takeaway to quote: *Do not let your production runtime dictate your iteration speed.* That single sentence, if internalized, will save you countless hours of waiting on pipelines that should have been split long ago. As you explore this pattern, keep an eye on how your evaluation loop changes. The real test is not whether you can run a workflow in production, but how quickly you can fail, learn, and try again. That speed is the true measure of progress, and it is the detail worth watching as you adopt this approach.

From InfoQ

AI workflows have two needs that trade off directly. Running reliably in production requires persisting and distributing every step so it survives crashes, deploys, and restarts. But that same machinery is what makes runs too heavy for the fast, throwaway loop you need to check an LLM's output quality. The properties that buy durability are the ones that kill iteration speed.

Read the original at InfoQ