database

database on Beyond Market Intelligence: a running collection of 16 stories we have gathered and hand-picked because they are worth your time. Every post here touches on database in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around database, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Connecting My LangGraph AI Agent to Postgres
Towards Data Science

Connecting My LangGraph AI Agent to Postgres

Connecting your LangGraph AI agent to a Postgres database unlocks powerful capabilities for data-driven workflows. This post details how to establish that connection, offering clear guidance for both local development and cloud deployment. We’ll explore setting up the backend using Docker for streamlined local testing, and then outline strategies for scaling to the cloud. For those tackling complex enterprise workflows, consider the recent exploration of an 8B AI model mirroring Claude Opus—a relevant challenge in managing substantial data sets.

One Document Type, a Million Files: Structured Extraction into the SQL Table RAG Queries
Towards Data Science

One Document Type, a Million Files: Structured Extraction into the SQL Table RAG Queries

Unlock the power of your enterprise data with structured extraction. This guide, "One Document Type, a Million Files," details a streamlined approach to transforming unstructured documents into SQL tables optimized for Retrieval-Augmented Generation (RAG) queries. In just one hour with two people, extract six to ten key fields, leveraging signals to ensure data integrity and filter accuracy. Explore how this method empowers efficient data access and analysis—a critical step toward future-focused data management.

Recursive CTEs: SQL’s Hidden Graph Traversal Engine
Towards Data Science

Recursive CTEs: SQL’s Hidden Graph Traversal Engine

Unlock the power of SQL for graph-like data manipulation with Recursive Common Table Expressions (CTEs). This practical guide reveals how CTEs function as SQL’s hidden engine for traversing hierarchies, identifying routes, and detecting cycles—capabilities often overlooked. Discover how to calculate degrees of separation and efficiently analyze complex relational structures. For a deeper dive into the nuances of context management within these workflows, explore "AI Agents Don’t Need More Context — They Need Typed Context."

Beyond Embedded: How DuckDB v2.0 Shifts Architecture Toward Distributed Network Capabilities
InfoQ

Beyond Embedded: How DuckDB v2.0 Shifts Architecture Toward Distributed Network Capabilities

DuckDB v2.0, codenamed "Cyanoptera," represents a significant architectural shift, moving beyond embedded processing toward distributed network capabilities. This preview release, built on over 10,000 commits, introduces a client/server mode for network connections alongside key improvements in extension portability and data type handling. Performance is enhanced through asynchronous I/O and storage optimizations, setting the stage for a more scalable future. General availability is slated for fall 2026. For a deeper dive into related architectural considerations, explore "Mini book: Architecture as a Socio-Technical Craft."

Harper Argues Against the Multi-System Stack and Releases 5.2
InfoQ

Harper Argues Against the Multi-System Stack and Releases 5.2

Harper is challenging the status quo of multi-system architectures, advocating for a single-runtime database platform that unifies application code and data. Recent benchmarks demonstrate significantly improved performance on live, personalized-data workloads compared to Vercel-based stacks. Version 5.2 further solidifies this approach, introducing a new record cache and increased throughput per node. Discover how Harper’s streamlined architecture empowers data-driven applications—for context, explore our analysis of Next.js 16.3’s recent performance improvements.

Microsoft Excel | Help & Support with your Formula, Macro, and VBA problems | A Reddit Community

Simple data collection form?

For teams needing simple, offline data collection, moving beyond outdated Access databases is achievable—even without SQL expertise. Our platform empowers you to build user-friendly forms directly within Excel, complete with a streamlined UI and a submit button to minimize input errors. Secure data storage is paramount, and we offer solutions to meet GDPR requirements. Explore transforming your data management with Excel’s capabilities, similar to how users leverage "Excel + Power Query and Power Automate" for broader integrations.

Machine Learning

Pandas API for DuckDB, PostgreSQL & ClickHouse — keeping computation inside the database[P]

Introducing memFrame, an open-source DataFrame API designed to transform your data workflow. Instead of importing data into Python, memFrame compiles operations directly to SQL, enabling computation within databases like DuckDB, PostgreSQL, and ClickHouse. This approach empowers users to leverage the power of their databases for data inspection, cleaning, statistics, and more—all while minimizing data transfer. We’re releasing features incrementally, prioritizing stability and user feedback. Explore this innovative architecture, including its built-in multiagent capabilities for natural language interaction with your data.

Running SQL Concurrently Across Three Remote DuckDB Servers with Quack
Towards Data Science

Running SQL Concurrently Across Three Remote DuckDB Servers with Quack

Explore a novel approach to data processing with "Running SQL Concurrently Across Three Remote DuckDB Servers with Quack." This experiment demonstrates a practical application of remote SQL execution, empowering users to leverage distributed resources for enhanced performance. Discover how Quack facilitates this process, offering a streamlined solution for complex queries. For those interested in building applications that accumulate understanding, consider "Designing a Persistent Knowledge Layer That Refuses to Guess," which details a vendor-neutral blueprint for RAG systems.

AWS Introduces Native Vector Search for DynamoDB
InfoQ

AWS Introduces Native Vector Search for DynamoDB

DynamoDB now offers native vector search, a significant advancement for developers working with semantic data. This integrated capability eliminates the need for separate vector databases, enabling you to store embeddings directly alongside application data and execute approximate nearest-neighbor queries within DynamoDB. Filtered similarity searches and configurable indexes further optimize performance for complex workloads. Explore this transformative feature and discover how it streamlines AI-powered applications—a concept further detailed in our article, "AWS Open-Sources Dogwood."

Microsoft Excel | Help & Support with your Formula, Macro, and VBA problems | A Reddit Community

How to get a table to match the number of rows, and row order, of a parent table

Need to streamline your data management? When adding new expense types to your parent table, automatically populate corresponding rows in linked tables—saving valuable time. This approach ensures consistent data structure and eliminates manual row insertion. It’s a powerful way to maintain data integrity and workflow efficiency. Discover how to configure this feature, mirroring your parent table's growth across related sheets. For troubleshooting formula errors that might arise, see our article, "Pls help - Need to fix formula with Spill Error," for helpful guidance.

  Token-maxxing is dead. Agentic memory is what comes next.
VentureBeat

Token-maxxing is dead. Agentic memory is what comes next.

The industry’s brief fascination with token-maxxing highlighted a crucial architectural lesson: the context window is a scarce resource. Now, after roughly 60 years of database development and just 18 months of agentic AI, we’re seeing a clear convergence. The future of agentic development lies in robust memory systems—semantic-search-backed, access-controlled, and even human-curated—that save and efficiently reuse previously generated insights. This shift promises a more economical and scalable approach, moving beyond the limitations of token-maxxing and ushering in a new era of AI productivity.

Stripe Uses Graph Search and State Machines to Automate Database Remediation
InfoQ

Stripe Uses Graph Search and State Machines to Automate Database Remediation

Stripe’s engineering team has achieved significant automation in database incident recovery, demonstrating a powerful application of graph search and state machines. By modeling their global infrastructure as a graph, they’ve created a system that automatically computes and executes remediation plans. This innovative approach minimizes downtime and reduces manual intervention, representing a future-focused strategy for managing complex, distributed systems. For further insights into the challenges of scaling AI infrastructure, explore our recent presentation with Martin Spier on keeping ChatGPT fast.

Wiz Discloses CosmosEscape, and Practitioners Debate What Customers Could Have Done
InfoQ

Wiz Discloses CosmosEscape, and Practitioners Debate What Customers Could Have Done

Wiz Research has revealed CosmosEscape, a significant security vulnerability impacting Azure Cosmos DB. This chain allowed an attacker to escape the Gremlin sandbox and obtain a platform-wide key, granting full read and write access to every database. While Microsoft swiftly blocked the initial entry point, remediation took nearly two years. The incident has sparked debate among security practitioners regarding shared responsibility and the true cost of this rearchitecture.

The Medallion Data Architecture: An Introduction
Towards Data Science

The Medallion Data Architecture: An Introduction

Navigating modern data pipelines can feel complex, but the Medallion Data Architecture offers a clear, practical framework. This guide introduces the Bronze, Silver, and Gold layers—a proven approach to structuring data for reliability and analytical readiness. We’ll explore each tier with a working Python and DuckDB example, empowering you to build robust data workflows. For a deeper dive into related challenges in AI agent memory management, see "Asana's AI agents share memory across your company — but not your secrets."

Bending Spoons to buy Airtable for $1.28B
TechCrunch

Bending Spoons to buy Airtable for $1.28B

Bending Spoons, the company behind popular apps like TikTok's algorithm, has acquired Airtable for $1.28 billion. Once valued at over $11 billion in 2021, Airtable’s valuation has adjusted, with recent secondary market trades around $4 billion. This acquisition signals a significant shift in the data management landscape, as Bending Spoons aims to empower users with accessible and innovative spreadsheet technology. For further insight into the evolving landscape of AI assistants, explore our recent analysis on "Apple finally fixed Siri."

Podcast: Rethinking Data: Moving From the Traditional Three-Tier Web Stack to Client-Side Event Sourcing
InfoQ

Podcast: Rethinking Data: Moving From the Traditional Three-Tier Web Stack to Client-Side Event Sourcing

Johannes Schickling challenges conventional wisdom in our latest podcast, "Rethinking Data." He details his journey moving beyond the traditional three-tier web stack to a local-first architecture, sharing his experience building Overtone—a music curation app—with client-side event sourcing and SQLite. This episode unpacks the practical trade-offs inherent in event sourcing and CRDTs, offering valuable insights for developers seeking a more agile data management approach. For further exploration of evolving architectures, see our article, "An Evolutionary Architecture Pattern for Managing AI’s Pace of Change."