observability

observability on Beyond Market Intelligence: a running collection of 12 stories we have gathered and hand-picked because they are worth your time. Every post here touches on observability in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around observability, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Sequoia-incubated Empirik launches with $21M to predict outages before they happen
TechCrunch

Sequoia-incubated Empirik launches with $21M to predict outages before they happen

Empirik, a Sequoia-incubated startup, emerges with $21 million in funding to redefine IT infrastructure management. Their mission: predict outages before they impact operations, mirroring Cursor's transformative approach to software engineering. This innovative platform empowers teams to proactively address potential issues, minimizing downtime and maximizing efficiency. Empirik’s predictive capabilities represent a significant advancement in data-driven infrastructure oversight. For deeper insights into related data trends, explore our article on "A group funded by Andreessen, Horowitz, and Brockman plans data center ads to sway midterms."

Presentation: Can Claude Fix Itself? Using LLMs for Incident Response
InfoQ

Presentation: Can Claude Fix Itself? Using LLMs for Incident Response

Incident response demands speed and precision. Join Anthropic reliability engineer Alex Palcuie as he shares practical lessons on leveraging Large Language Models (LLMs) for real-world troubleshooting. This presentation clarifies where AI excels—acting as a superhuman observer of logs and traces—while also highlighting persistent challenges in root-cause analysis, specifically distinguishing causation from correlation. Palcuie outlines how engineering leaders can effectively integrate AI into on-call workflows, preserving crucial human expertise.

Microsoft Moves AI Governance From Policy to Runtime Enforcement
InfoQ

Microsoft Moves AI Governance From Policy to Runtime Enforcement

Microsoft is reshaping AI governance, moving beyond policy creation to runtime enforcement. Their new architecture, spanning nine domains and four core functions—policy, control, visibility, and proof—directly links governance requirements with real-world application operation. This approach ensures continuous evaluation, observability, and robust audit trails, empowering organizations to confidently verify AI compliance. As enterprises increasingly leverage AI agents, understanding this shift is critical; consider “Enterprises winning with AI agents are limiting how much the agents can do alone” for further insights.

VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push
VentureBeat

VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push

VentureBeat significantly expands its enterprise AI research capabilities with the appointment of Rob Strechay as its first Lead Analyst. Strechay, formerly of theCUBE Research, brings three decades of experience across practitioner, executive, and analyst roles, uniquely positioning him to address the critical data needs of technical decision-makers. His focus will initially encompass cloud infrastructure, data infrastructure, and AI security, complementing VentureBeat’s VB Pulse surveys—including recent findings on agentic orchestration—to provide objective insights for navigating the evolving AI landscape.

Cloudflare Adds Agent Tracing, with Truncation Limits and Uneven Payload Defaults
InfoQ

Cloudflare Adds Agent Tracing, with Truncation Limits and Uneven Payload Defaults

Cloudflare has expanded its tracing capabilities with the introduction of Agent Tracing, now incorporating spans for agent invocations, model calls, tool runs, and approvals within Workers traces. This feature enables session replay, offering deeper insights into agent workflows. While traces provide valuable context, users should note that they are not lossless, and payloads may be truncated due to default settings that vary by framework. Starting October 1, 2026, each span will be a billable event.

Presentation: Microservices Platforms: When Team Topologies Meets Microservices Patterns
InfoQ

Presentation: Microservices Platforms: When Team Topologies Meets Microservices Patterns

Accelerate your microservices delivery with a strategic blend of Team Topologies and proven patterns. Chris Richardson’s presentation explores how internal platforms, built around six key areas—security, observability, build, and deployment—can minimize cognitive load for development teams. Richardson shares practical strategies to avoid common platform engineering challenges and maximize efficiency. Discover how to empower stream-aligned teams and unlock faster innovation. For a deeper dive into the broader context, see our related article, "Platform Engineering Maturity Emerges as a Key Differentiator for Enterprise AI Success."

How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud
VentureBeat

How is your enterprise tracking AI agent telemetry? Groundcover thinks it should never leave your cloud

The rise of AI agents is fundamentally reshaping enterprise data management, particularly how telemetry is tracked. Groundcover thinks it should never leave your cloud, offering a compelling alternative to traditional observability platforms. With $160 million in funding, the company is challenging established players like Datadog and Splunk by prioritizing customer-controlled data storage and a predictable, host-based pricing model. Explore how this approach, combined with eBPF technology, is transforming observability into infrastructure for autonomous software, as discussed further in our recent article, "Smallest.

Enterprise AI agents can't talk to each other, can't be trusted with permissions, and can't be audited — 5 startups are already fixing that
VentureBeat

Enterprise AI agents can't talk to each other, can't be trusted with permissions, and can't be audited — 5 startups are already fixing that

Enterprise AI agents promise transformative work capabilities, but a crucial infrastructure gap remains: ensuring secure communication, reliable authorization, and comprehensive auditing. Five innovative startups are addressing this challenge, focusing on orchestration, observability, connectivity, and security. From BAND’s coordination layer to Arcade's secure runtime, these solutions are laying the groundwork for a future where AI agents collaborate seamlessly and securely. As Meta envisions billions of personal AI agents within five years, this foundational work is increasingly vital.

Target SVP says its real AI moat isn't the models — it's everything built around them
VentureBeat

Target SVP says its real AI moat isn't the models — it's everything built around them

Target SVP Siobhán McFeeney asserts that Target’s competitive advantage in AI isn’t solely reliant on advanced models, but rather the robust infrastructure built around them. The company’s approach prioritizes deliberate agent deployment, ensuring they address high-value problems and “earn” autonomy through demonstrable results. This framework, encompassing architecture, taxonomy, and rigorous observability, enables scalable AI investment and allows Target to strategically leverage models—from frontier to specialized—for optimal cost-benefit. For deeper insight into agent architecture, explore Microsoft’s recent reference architecture for AI agents on AKS.

Grafana Assistant Expands to More Than 30 Data Sources
InfoQ

Grafana Assistant Expands to More Than 30 Data Sources

Grafana Assistant now empowers users to explore observability insights across a broader landscape, integrating with more than 30 diverse data sources. This expansion allows for natural language queries and correlations, streamlining data analysis and accelerating troubleshooting. Leverage AI to transform how you understand your systems, moving beyond siloed views. For a deeper dive into related AI projects, see our recent article, "Recent project I worked on: End to End Edge ML platform," demonstrating practical applications of AI-driven solutions.

Expedia Uses AI Driven Service Telemetry Analyzer to Accelerate Incident Investigation
InfoQ

Expedia Uses AI Driven Service Telemetry Analyzer to Accelerate Incident Investigation

Expedia Group is accelerating incident investigation with STAR, a novel AI-assisted observability platform. Built on FastAPI, Datadog, and other key technologies, STAR leverages LLMs to analyze service telemetry and generate root cause assessments, streamlining workflows for engineers. This innovative approach keeps engineers informed while significantly reducing resolution times. STAR represents a future-focused evolution in production incident management, demonstrating how AI can empower data-driven response. For deeper insights into production AI, explore our coverage of QCon AI New York 2026.

Linkerd 2.20 Delivers Smarter Traffic Management and Dramatic Efficiency Gains
InfoQ

Linkerd 2.20 Delivers Smarter Traffic Management and Dramatic Efficiency Gains

Linkerd 2.20 significantly elevates Kubernetes networking with smarter traffic management and dramatic efficiency gains. This release, announced by the Linkerd community, delivers key enhancements across performance, observability, and control. As a CNCF-graduated service mesh, Linkerd remains the leading lightweight choice for Kubernetes, empowering teams to optimize application delivery. Explore the new features to discover how Linkerd 2.20 streamlines operations and unlocks greater resource utilization within your existing infrastructure.