Uber's pointer-based federation brings zero-downtime scale to 16K datasets.

Uber has successfully decentralized its Hive data warehouse, migrating an impressive 16,000 datasets totaling over 10 petabytes through pointer-based federation.

3 min readInfoQ
Uber's pointer-based federation brings zero-downtime scale to 16K datasets.

Uber's migration of 16,000 datasets across more than 10 petabytes is the kind of infrastructure story that usually stays buried in engineering blogs, but it deserves attention. Pointer-based federation is not just a clever trick for moving data without downtime. It is a practical answer to a problem every large organization eventually hits: how do you decentralize ownership without breaking the systems people already rely on? Uber's approach proves that you can have domain-specific datasets and strict governance at the same time, and that is a meaningful step forward for anyone wrestling with data sprawl.

For readers who manage data platforms, the key takeaway is the zero-downtime piece. Migrating data is typically a high-stakes operation where the risk of service disruption looms over every decision. Uber's use of pointer-based federation flips that script by decoupling the logical view of the data from its physical location. That means teams can move datasets, reassign ownership, and rebalance workloads without forcing users to change their queries or endure maintenance windows. The practical implication is straightforward: you do not have to choose between modernizing your architecture and keeping the lights on. That is a trade-off most platforms have accepted as unavoidable, and it is refreshing to see it challenged with results.

The other piece worth highlighting is the emphasis on strict ACL enforcement and governance. Decentralization often gets treated as a synonym for chaos, where domain teams gain autonomy at the expense of oversight. Uber's migration suggests otherwise. By keeping access control lists intact and enforcing them consistently across federated datasets, they show that distributed ownership does not have to mean diluted security. For data leaders, this is a useful counterpoint to the fear that moving toward domain-specific architectures will introduce compliance risks. The governance layer is not an afterthought here; it is baked into the migration strategy itself.

What stands out most is that this was not a greenfield project. Uber migrated an existing Hive warehouse with years of accumulated data and workflows. That is the situation most enterprises are actually in, regardless of how much they talk about cloud-native transformations. Pointer-based federation offers a path that respects the past while enabling a more scalable future. If you are staring down a similar migration, the lesson is not that you need a new platform or a rewrite of your data model. It is that the connective tissue between systems matters just as much as the systems themselves. That is a concrete, actionable insight, and it is one worth carrying into your next architecture review.

From InfoQ

Uber has decentralized its Hive data warehouse, migrating 16,000 datasets totaling over 10 petabytes using pointer-based federation. The migration ensures zero downtime, strict ACL enforcement, improved governance, and scalable, domain-specific datasets for analytics and machine learning workloads.

Read the original at InfoQ