Spanner Omni

Google's Spanner Omni goes live, swapping hardware clocks and storage for software

Google made Spanner Omni generally available, swapping its own hardware clocks and storage for software abstractions.

4 min readInfoQ
Google's Spanner Omni goes live, swapping hardware clocks and storage for software

Google just made its distributed SQL database available everywhere: on-premises, across clouds, and on a laptop. But the real story is what Google had to give up to get there. Spanner Omni swaps Colossus for a Colossus-like abstraction layer and replaces TrueTime with software-based time synchronization. That trade is the entire ballgame, and it deserves more scrutiny than the general availability announcement is getting. The absence of an availability SLA tells you everything about the trade-offs. Spanner's original magic was that it leaned on specialized hardware and tightly coupled storage to deliver global consistency with low latency. Omni keeps the SQL semantics and the replication logic, but it runs on whatever storage you already have and syncs time through software instead of atomic clocks. That is a meaningful engineering compromise, and practitioners are right to price it in tail latency and operational toil. You are not getting the same Spanner; you are getting a Spanner that has traded some guarantees for portability. That is a fair trade, but only if you know what you are giving up. The absence of an availability SLA is not a footnote. It is the headline. This move echoes a broader tension we have been tracking in our coverage. When Cloudflare resolves data leak risk across container boundaries, it is because the abstraction between hardware and software created a subtle trust gap. Spanner Omni faces the same pressure. By swapping Colossus for a Colossus-like abstraction layer and replacing TrueTime with software-based time synchronization, Google has traded physical guarantees for logical ones. That is a fair trade for portability, but it changes the operational calculus. Practitioners are already pricing this move in tail latency and operational toil, which is the honest way to evaluate it. There is no availability SLA here, and that absence is not a footnote. It is the point. The decision to decouple Spanner from Google's custom hardware is a pragmatic acknowledgment that the future of data infrastructure is heterogeneous. But pragmatism comes with a tax. TrueTime was built on specialized hardware clocks because distributed consensus demands precise time; replacing it with software-based synchronization means your mileage will vary based on network conditions, clock drift, and the kindness of your cloud provider's hypervisor. Similarly, swapping Colossus for a Colossus-like abstraction layer gives you the portability without the performance guarantees. This is not a criticism of the engineering, which is genuinely impressive, but a warning about expectations. You are not getting Google's internal Spanner. You are getting a flexible, software-defined approximation of it. And as we have seen with Cloudflare resolves data leak risk across container boundaries, abstraction layers can introduce their own edge cases and failure modes that are easy to underestimate. The absence of an availability SLA is the tell. Google is effectively saying: we will give you the software, but not the guarantee that comes with it. That is a reasonable trade for a laptop demo or a dev environment, but it changes the calculus for production workloads. Practitioners are already pricing this in tail latency and operational toil, which is another way of saying the abstraction is not free. You trade hardware clocks for software time sync, and you trade Colossus for something that looks like it but carries its own assumptions. This is the same tension that runs through the broader conversation about distributed systems: every layer of abstraction solves a problem and introduces new ones. The move also echoes a pattern we saw when Cloudflare disclosed a cross-tenant data exposure vulnerability in containers and sandboxes, where the abstraction that makes infrastructure portable also makes its boundaries harder to reason about. Cloudflare resolves data leak risk across container boundaries. The convenience of running the same database engine anywhere comes with a new class of operational questions that the hardware era never had to ask. The absence of an availability SLA is the tell. Google is not promising five nines here; it is promising a consistent experience and letting you price the tail latency yourself. That is a meaningful shift for teams that treat SLAs as a safety net.

From InfoQ

Google has made Spanner Omni generally available, running its distributed SQL database on-premises, across clouds or on a laptop. Doing so meant replacing Colossus with a Colossus-like abstraction layer and TrueTime with software-based time synchronization. There is no availability SLA, and practitioners are pricing the move in tail latency and operational toil.

Read the original at InfoQ