Postgres

Postgres: Your Durable Orchestrator for Workflows Without External Tools

Most teams assume durable workflows demand an external orchestrator, but Postgres can handle that heavy lifting on its own.

4 min readInfoQ
Postgres: Your Durable Orchestrator for Workflows Without External Tools

The appeal of dropping a dedicated orchestrator in favor of Postgres is not about being contrarian; it is about recognizing that your database already does the heavy lifting. Raman Varma's article on implementing durable workflows without an external tool is a pragmatic counterpoint to the reflex of reaching for another service the moment a job needs to survive a crash. We have seen teams bolt on orchestration layers for what are essentially transactional state machines, and the complexity budget gets spent before a single business rule is written. This approach suggests a leaner path: use what you have, enforce idempotency with primary keys, and let `SKIP LOCKED` handle the concurrency that would otherwise require a dedicated queue.

This matters because the operational cost of running a separate orchestrator is real, even if it is often hidden. Every dependency you add is a system you must keep alive, patch, and reason about during an incident. Varma's method of persisting workflow sleeps and human approvals as database state is particularly interesting, because it turns what is usually an external concern into a transaction. If a workflow is waiting on a person to click a button, that state can live in a row, and it will survive a restart without a separate recovery process. That is the kind of resilience that feels obvious in hindsight, but it requires the discipline to model time and external input as data. For teams already fluent in SQL, this is not a step backward; it is a way to reduce the operational surface area while keeping the guarantees that matter. Related work on scaling concurrent workloads, like the Scale Sandboxes Instantly: A New Approach to Concurrent AI Workloads piece from Modal's engineers, shows a similar appetite for shedding unnecessary infrastructure, even if their problem domain is different. The instinct to simplify is shared, even when the systems look nothing alike.

Our take is that this works best when your workflow is essentially a series of database operations with a few human checkpoints. The moment you need complex retries with exponential backoff, or you are fanning out thousands of parallel jobs with dynamic task graphs, you will start reimplementing a framework inside your database. That is a trade-off, and it is worth naming. You are trading the flexibility of a purpose-built tool for the comfort of a single source of truth. What you gain is a system that is easier to debug because the state is visible in a table, and easier to operate because there is no second cluster to worry about. We would tell a reader to look at their current pain points: if the orchestrator is mostly idle and the real bottleneck is coordinating a few long-running steps, this design is likely a better fit than yet another microservice. Removing state where you can simplifies operations, and the same logic applies to workflow coordination, as shown in Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol.

The practical takeaway is a question: can you express your workflow as a set of checkpoints in a transaction? If yes, the orchestration layer might be the most expensive thing you do not need. The specific detail to watch is how you handle leases during long crashes, because that is where the subtle bugs live. A row that is locked forever is a workflow that never completes, and no amount of clever SQL will save you from a missing timeout. That is the edge of this pattern, and it is where your team will either find elegance or frustration. We would suggest starting with a single critical process, not an entire platform, and measuring whether the operational burden actually shrinks. That is the real test, and it is the one that will tell you if Postgres alone is enough.

From InfoQ

Postgres can serve as the durable state store and coordination layer for workflows, eliminating the need for an external orchestrator. SKIP LOCKED enables concurrent work processing, primary-key checkpoints enforce idempotency, and leases support crash recovery. Workflow sleeps and human approvals can also be persisted as database state and survive restarts.

Read the original at InfoQ