workflow automation
workflow automation on Beyond Market Intelligence: a running collection of 260 stories we have gathered and hand-picked because they are worth your time. Every post here touches on workflow automation in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around workflow automation, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Forward-deployed engineering is how enterprise AI learns
Forward-deployed engineering (FDE) is rapidly reshaping enterprise AI, but its true value isn't always clear. Zeta’s Neej Gore unpacks the nuances, distinguishing between FDE that builds lasting product advantage and that which simply accumulates delivery labor. The test? Does each subsequent deployment leverage more product and fewer unknowns? This piece explores how to evaluate FDE, track its impact, and ensure it fuels a system of intelligence – ultimately, a product that gets better at understanding.

Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads
Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the latest iterations of its powerful large language models, alongside a significant 75% cost reduction for Fable cache reads. These models prioritize sustained problem-solving, demonstrating substantial improvements on benchmarks like Terminal-Bench and AutomationBench. Crucially, Anthropic is also introducing Enterprise Frontier Safeguards (EFS), allowing organizations to retain monitoring data within their own infrastructure. This release addresses evolving enterprise needs for capable, economical, and governable AI agents—a shift underscored by recent cybersecurity evaluations.

AI is redefining the workforce — and most planning models aren’t ready
AI is rapidly reshaping the workforce, and traditional planning models are struggling to keep pace. Fragmented data across HR, finance, and procurement leaves executives blind to how workforce decisions impact business outcomes. Recent SAP research reveals a significant gap: while organizations plan for AI's impact on productivity, few address its influence on job design and organizational structure. To navigate this shift, explore SAP Workforce Planning and SuccessFactors innovations for a clearer view of work's value.

OpenClaw 2.0 is here, ushering in the era of 'multiplayer' AI coding: What it means for enterprises
OpenClaw 2.0 is here, marking a significant shift toward enterprise-ready AI coding. Building on the viral momentum of earlier versions, this update transforms OpenClaw from a personal agent harness into a collaborative platform designed for teams and shared infrastructure. Key additions include a rebuilt browser interface, shared cloud sessions, and enhanced security features like role-based permissions and auditing. For organizations, OpenClaw 2.0 envisions agents as a shared operational layer, not just individual developer tools—a concept DoorDash recently explored with its Flux platform.

DoorDash’s Flux Runs 130,000 Engineering Tasks Through Cloud-Based Agents
DoorDash has significantly streamlined its engineering workflows by migrating 130,000 tasks monthly to its proprietary Flux cloud platform. This transition moves agent workloads off developer laptops and leverages isolated Firecracker microVMs for secure, automated operations. Flux now facilitates over 25,000 automated code reviews weekly, demonstrating a clear shift toward future-focused data management. For deeper insight into the evolving landscape of AI agents and their infrastructure needs, explore "AI agents need their own identity before they need a gateway."

AI agents need their own identity before they need a gateway
Enterprise AI has entered a new era, moving beyond simple assistants to autonomous agents capable of complex workflows. This shift introduces a fundamental security challenge: authentication confirms identity, but it doesn't guarantee ongoing trust. Traditional security controls offer limited visibility into an agent’s actions after authentication, creating new runtime risks like goal drift and memory poisoning. To address this, organizations must embrace runtime trust – continuously validating AI behavior and ensuring alignment with organizational policy.

Orchestration is the new challenge for CX in the age of AI agents
The rise of AI agents presents a new challenge for customer experience: orchestration. As enterprises rapidly deploy AI across channels, many are struggling to integrate these tools with legacy systems, creating fragmented customer journeys and overburdened human agents. Tata Communications’ Gaurav Anand explains that the shift is moving away from simple automation toward intelligent orchestration—connecting tasks and delivering end-to-end outcomes with a shared understanding of the customer. Discover how this approach, underpinned by a common enterprise ontology, can transform CX.

Cloudflare OS: Cloudflare's Open-Source Corporate AI Platform Built on a Capability-Based Model
Cloudflare OS, now open-source, represents a progressive shift in enterprise AI. This capability-based platform empowers teams to generate work artifacts rooted in company knowledge, automate workflows with optimized efficiency, and build customized work software within a secure environment. Unlike approaches prioritizing full autonomy, Cloudflare OS strategically employs AI assistance only when needed, maximizing cost-effectiveness. See how Cloudflare itself leveraged this approach to significantly reduce GitHub issues, as detailed in "Cloudflare Cuts Astro Github Issues by 85% with AI Agents."

Enterprises winning with AI agents are limiting how much the agents can do alone
Enterprises are discovering a critical truth about AI agents: unrestrained autonomy isn't synonymous with superior performance. While the initial focus was on maximizing agent independence, current deployments reveal that controlled, narrowly-scoped agents, coupled with strategic human checkpoints, are proving far more sustainable. Gartner forecasts that over 40% of agentic AI projects won't reach 2028, highlighting a widening gap between capability and responsible AI maturity.

Cloudflare Cuts Astro Github Issues by 85% with AI Agents
Cloudflare significantly enhanced developer productivity by leveraging AI agents to manage GitHub issues, achieving an 85% reduction in processing time. This innovative application of agentic AI within GitHub Actions streamlines issue triage, automating workflows and accelerating software engineering cycles. Utilizing Cloudflare Workers and Flue, the system incorporates a “human-in-the-loop” approach, ensuring quality while maximizing efficiency.

Slack wants to drag AI coding out of the terminal and into the group chat
Slack is redefining AI coding, moving it from isolated terminals into collaborative group chats with Slack Code. This new product embeds AI coding agents – including Anthropic's Claude Code, Cognition's Devin, and others – directly into dedicated Slack channels, enabling teams to collectively review, steer, and ship software. Slack is betting that multiplayer AI, where everyone can see and contribute, will lead to higher quality and faster development, effectively removing code as the primary bottleneck.

NanoClaw comes to Slack, letting you create persistent AI agent teams and colleagues from a single message
NanoCo is simplifying the integration of AI agents into Slack with its new NanoClaw Slack integration, enabling users to create persistent teams of AI colleagues from a single message. Unlike previous attempts at AI integration that often felt clunky, NanoClaw allows for the effortless creation of specialized agents, each with custom skills, workflows, and even avatars.

Serval’s super agent Catalyst creates roving background agents to identify and fix IT issues before they’re ticketed
Serval is making its AI agent, Catalyst, generally available Thursday, empowering teams to automate enterprise workflows with unprecedented ease. This "super agent" analyzes ticket history, SOPs, and instructions to draft workflows, skills, and dashboards – even proactively identifying and fixing IT issues before they reach a ticket queue. Unlike competitors, Catalyst operates as a single administrative layer, moving from opportunity discovery to deploying proactive agents.

TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Claude Managed Agents
TrueFoundry introduces TrueForge, a new open-source AI agent harness designed to empower enterprise developers and reduce costs. Built by former Meta and Google engineers, TrueForge offers a vendor-neutral solution, compatible with various AI models and deployable across different infrastructures. Initial testing reveals impressive cost savings—up to 75% less than Anthropic’s Claude Managed Agents—achieved through intelligent context engineering.

Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation history locally
Block, the technology company behind Square and Cash App, is open-sourcing Berd, a desktop application designed to streamline AI agent workflows. Available now on GitHub under the Apache 2.0 license, Berd provides a unified workspace for users to manage projects, models, and conversation history locally. This innovative tool empowers users to explore AI capabilities across various agents and harnesses, offering a future-focused alternative to fragmented experiences.

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
Recent VentureBeat research reveals a concerning trend: 85% of companies that experienced an AI mistake are accelerating their move toward automated deployments, even as trust in automated evaluation rises. While automated checks are gaining traction, nearly half of surveyed enterprises still see test-approved AI features disappoint customers. This shift highlights a growing gap between evaluation confidence and real-world outcomes, prompting many to prioritize anomaly detection and issue resolution, as evidenced by the surging demand for platforms like Raindrop.ai.

How Heidi built production-ready AI for healthcare at global scale
Building production-ready AI for healthcare at scale demands a robust architecture, particularly when navigating stringent compliance requirements. Australian AI Care Partner, Heidi, provides a compelling case study. Its AI Scribe automates administrative tasks for clinicians across 190 countries, processing roughly 2.7 million patient interactions weekly. This global reach is underpinned by a data-first approach, leveraging MongoDB Atlas for flexible data management and AI-ready features like Vector Search. As Heidi’s co-founder, Yu Liu, emphasizes, "Reliability engineering is trust engineering.”

Grab Cuts Mechanical Analytics Work From 44% to 30% with AI Agents
Grab has demonstrably transformed its analytics workflows with AI agents, achieving a significant 30% reduction in mechanical analyst work since February – a 44% decrease. This progress stems from a powerful combination of agent autonomy, certified data, contextual awareness, and crucial human oversight. Self-service analytics are increasingly handling routine metric, data, and SQL requests, freeing analysts for higher-value tasks. Interested in the underlying architectural principles? Explore "Agentic Fitness Functions" for a deeper dive into extending evolutionary architecture.

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek’s V4 Flash, initially lauded as a "total monster" for its impressive leaderboard performance and remarkably low pricing, is experiencing a shift in perception. Recent testing reveals it completes only 53.8% of complex agent tasks in real-world scenarios. Simultaneously, DeepSeek is adjusting its pricing model, increasing rates by as much as 1,100% for certain token types.

Meta Open-Sources Muse Glimmer: A 30B Local Agentic Model Optimised for On-Device Execution
Meta AI Research has unveiled Muse Glimmer, a significant advancement in on-device AI. This 30-billion-parameter, open-weight model, released under the Apache 2.0 license, empowers autonomous agents and complex task execution directly on consumer GPUs—eliminating the need for cloud dependencies. Utilizing a multi-stage training process, Glimmer delivers efficient performance and supports multimodal inputs, streamlining coding and automation. Explore this future-focused solution, and discover how it transforms local workflows; for broader context on enterprise AI initiatives, see our related article on IBM’s partnership with OpenAI.

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Google is accelerating AI innovation with the release of Gemini 3.7 Flash, its "most intelligent workhorse model yet" for coding and agentic workflows. This upgrade prioritizes diligent planning and disciplined execution, showing significant gains in debugging, web development, and enterprise automation—potentially reducing human intervention. Notably, Google is offering a 50% introductory price cut through the end of 2026, making it a compelling option for high-volume applications.

Why Capital One built its multi-agent AI platform around open-weight models
At VB Transform 2026, Capital One’s Kel Vanee detailed the bank’s strategic shift toward building AI, not just using it. Capital One constructed a scalable, multi-agent AI platform centered around deeply customized open-weight models, leveraging proprietary data for enhanced accuracy and extensibility. This approach, underpinned by prior investments in data transformation and cloud adoption, enables the bank to optimize workflows, from fraud detection to customer service, and even automate internal infrastructure tuning.

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges
Writer today unveiled Palmyra X6, a new AI agent model poised to significantly reduce costs for enterprise users. Paired with its rebuilt agent orchestration “harness,” Palmyra X6 delivers an average of 52% lower operational costs, alongside a 48% speed improvement and 10% quality boost. Leveraging a post-trained version of GLM-5.2, Writer emphasizes control and cost transparency, offering governance tools and multi-model support—a strategy echoing the shift towards pragmatic AI adoption, as explored in "Why Capital One built its multi-agent AI platform around open-weight models."

Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least
Confidence in automated agent evaluation surged this July, nearly tripling to 13% across 108 enterprises – a shift largely driven by those yet to experience a “false-confidence” failure. Critically, the failure rate of agents passing evaluations but then causing customer issues remained unchanged at just under half. While trust is rising, enterprises are simultaneously increasing investment in human review workflows, hedging against evaluations that don’t always reflect real-world outcomes.