1 min readfrom InfoQ

Cloudflare Details Unified Data Platform Where Billing Workloads Account for 53% of Queries

Our take

Cloudflare’s Town Lake, an internal unified data platform, represents a significant advancement in data management. Processing approximately 91,000 queries, with billing workloads accounting for 53% of usage, Town Lake consolidates operational, billing, security, and business data through an innovative lakehouse architecture. Built with Trino, Iceberg, R2, and DataHub, it empowers governed cross-system analytics and natural language access via Skipper, an AI analytics agent.
Cloudflare Details Unified Data Platform Where Billing Workloads Account for 53% of Queries

Cloudflare’s unveiling of Town Lake and Skipper represents a significant, albeit quietly impactful, step in the evolution of data platform architecture, particularly for organizations grappling with the complexities of modern, distributed systems. The sheer scale of the platform – handling approximately 91,000 queries, with billing workloads dominating at 53% – underscores the growing need for unified data access and analysis. This isn’t just about consolidating data; it’s about enabling cross-functional insights that were previously siloed, a challenge also highlighted by Cycle's recent introduction of an EU Control Plane as Sovereignty Debate Continues demonstrating the increasing demands around data governance and regional compliance. Furthermore, the limitations surrounding Claude’s availability on Microsoft Foundry, as noted in Claude Reaches GA on Microsoft Foundry: European Enterprises Cannot Deploy It, illustrate the difficulties organizations face when trying to seamlessly integrate AI-powered analytics into existing infrastructure. Cloudflare’s approach, leveraging a lakehouse architecture built on open-source components like Trino, Iceberg, and DataHub, offers a compelling alternative to proprietary solutions, and potentially a more adaptable one.

The decision to build Town Lake around a lakehouse architecture is particularly noteworthy. While data warehouses have traditionally been the go-to solution for structured data, they often struggle to accommodate the variety and velocity of data generated by modern applications. Lakehouses, combining the flexibility of data lakes with the governance and performance of data warehouses, are rapidly emerging as the preferred architecture for organizations seeking to unlock the full potential of their data. Cloudflare’s adoption of Iceberg for data management signals a commitment to a robust and evolving standard, minimizing vendor lock-in and fostering interoperability. Integrating R2 for object storage adds another layer of flexibility, allowing them to leverage cost-effective storage for less frequently accessed data. The inclusion of DataHub for data discovery and metadata management is crucial, ensuring that users can easily find and understand the data available to them, a common pain point in large, complex data environments. This focus on open standards and modular components reflects a pragmatic and future-focused approach to data management.

What’s perhaps most compelling is the integration of Skipper, the AI analytics agent. Natural language access to data is no longer a futuristic fantasy; it's becoming a practical reality. Empowering users to query data using plain language significantly lowers the barrier to entry for data analysis, enabling a wider range of stakeholders to leverage data-driven insights. While AWS’s recent introduction of Amazon S3 Annotations speaks to improving data context and searchability, Skipper takes it a step further by facilitating direct data exploration and analysis through natural language. This underscores a broader trend toward democratizing data access and making it more intuitive for non-technical users, ultimately driving greater productivity and innovation. Cloudflare’s focusing on billing queries initially also hints at a strategic prioritization of areas with immediate ROI and a clear need for data-driven optimization.

Ultimately, Cloudflare’s Town Lake and Skipper represent a model for how organizations can build scalable, governed, and AI-powered data platforms using open-source technologies. It demonstrates that a unified data platform doesn’t have to be a monolithic, proprietary system. The emphasis on accessibility and natural language querying, combined with the lakehouse architecture, positions Cloudflare to unlock significant value from its data, and provides a blueprint for other organizations seeking to modernize their data infrastructure. The key question now is whether this approach—favoring open standards and modularity—will become the dominant pattern for building enterprise-grade data platforms, or if proprietary solutions will continue to hold sway.

Cloudflare details Town Lake, an internal unified data platform, and Skipper, an AI analytics agent unifying access to operational, billing, security, and business data. The platform processed ~91K billing queries, with billing forming majority usage. Built on a lakehouse architecture using Trino, Iceberg, R2, and DataHub, it enables governed cross-system analytics and natural language access.

By Leela Kumili

Read on the original site

Open the publisher's page for the full experience

View original article