generative AI automation
generative AI automation on Beyond Market Intelligence: a running collection of 278 stories we have gathered and hand-picked because they are worth your time. Every post here touches on generative ai automation in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around generative ai automation, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities
Google continues to advance its AI capabilities with the release of Gemini 3.8 Flash, offering distinct models tailored for specific needs. The standard 3.8 Flash excels at agentic tasks and software development, demonstrating significant performance improvements over its predecessor and rivaling larger models at a reduced cost. Notably, Flash Cyber represents a substantial leap in cybersecurity, autonomously identifying and patching vulnerabilities with impressive efficiency—already securing Google's own code. For those exploring enterprise AI, consider “Forward-deployed engineering is how enterprise AI learns” for deeper insights.

Forward-deployed engineering is how enterprise AI learns
Forward-deployed engineering (FDE) is rapidly reshaping enterprise AI, but its true value isn't always clear. Zeta’s Neej Gore unpacks the nuances, distinguishing between FDE that builds lasting product advantage and that which simply accumulates delivery labor. The test? Does each subsequent deployment leverage more product and fewer unknowns? This piece explores how to evaluate FDE, track its impact, and ensure it fuels a system of intelligence – ultimately, a product that gets better at understanding.

Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads
Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the latest iterations of its powerful large language models, alongside a significant 75% cost reduction for Fable cache reads. These models prioritize sustained problem-solving, demonstrating substantial improvements on benchmarks like Terminal-Bench and AutomationBench. Crucially, Anthropic is also introducing Enterprise Frontier Safeguards (EFS), allowing organizations to retain monitoring data within their own infrastructure. This release addresses evolving enterprise needs for capable, economical, and governable AI agents—a shift underscored by recent cybersecurity evaluations.

OpenClaw 2.0 is here, ushering in the era of 'multiplayer' AI coding: What it means for enterprises
OpenClaw 2.0 is here, marking a significant shift toward enterprise-ready AI coding. Building on the viral momentum of earlier versions, this update transforms OpenClaw from a personal agent harness into a collaborative platform designed for teams and shared infrastructure. Key additions include a rebuilt browser interface, shared cloud sessions, and enhanced security features like role-based permissions and auditing. For organizations, OpenClaw 2.0 envisions agents as a shared operational layer, not just individual developer tools—a concept DoorDash recently explored with its Flux platform.

AI agents that pass authentication can still drift, expose data, or get memory-poisoned
Securing AI agents requires a shift in perspective. While gateways are often the initial defense, they're frequently deployed before foundational identity and attribution layers are in place, creating a significant vulnerability. Recent events, like the CISA advisory regarding a LiteLLM flaw, highlight this risk. Prioritize establishing agent inventory, distinct identities, and task-scoped credentials *before* relying on runtime enforcement. Start with the basics – identifying and naming your agents – to build a robust security foundation.

AI agents need their own identity before they need a gateway
Enterprise AI has entered a new era, moving beyond simple assistants to autonomous agents capable of complex workflows. This shift introduces a fundamental security challenge: authentication confirms identity, but it doesn't guarantee ongoing trust. Traditional security controls offer limited visibility into an agent’s actions after authentication, creating new runtime risks like goal drift and memory poisoning. To address this, organizations must embrace runtime trust – continuously validating AI behavior and ensuring alignment with organizational policy.
py-evoFE: Automated Evolutionary Feature Engineering for Tabular ML in Python (Genetic Algorithms + Scikit-Learn + Polars) [P]
Announcing py-evoFE (v0.3.0), an open-source Python library designed to automate and optimize feature engineering for tabular machine learning. Leveraging genetic algorithms alongside Scikit-Learn and Polars, py-evoFE intelligently discovers and combines feature transformations, addressing a critical bottleneck in model development. Unlike brute-force methods that generate excessive, noisy features, py-evoFE employs evolutionary selection to produce compact, high-impact recipes.

Orchestration is the new challenge for CX in the age of AI agents
The rise of AI agents presents a new challenge for customer experience: orchestration. As enterprises rapidly deploy AI across channels, many are struggling to integrate these tools with legacy systems, creating fragmented customer journeys and overburdened human agents. Tata Communications’ Gaurav Anand explains that the shift is moving away from simple automation toward intelligent orchestration—connecting tasks and delivering end-to-end outcomes with a shared understanding of the customer. Discover how this approach, underpinned by a common enterprise ontology, can transform CX.

Enterprises winning with AI agents are limiting how much the agents can do alone
Enterprises are discovering a critical truth about AI agents: unrestrained autonomy isn't synonymous with superior performance. While the initial focus was on maximizing agent independence, current deployments reveal that controlled, narrowly-scoped agents, coupled with strategic human checkpoints, are proving far more sustainable. Gartner forecasts that over 40% of agentic AI projects won't reach 2028, highlighting a widening gap between capability and responsible AI maturity.

Slack wants to drag AI coding out of the terminal and into the group chat
Slack is redefining AI coding, moving it from isolated terminals into collaborative group chats with Slack Code. This new product embeds AI coding agents – including Anthropic's Claude Code, Cognition's Devin, and others – directly into dedicated Slack channels, enabling teams to collectively review, steer, and ship software. Slack is betting that multiplayer AI, where everyone can see and contribute, will lead to higher quality and faster development, effectively removing code as the primary bottleneck.

NanoClaw comes to Slack, letting you create persistent AI agent teams and colleagues from a single message
NanoCo is simplifying the integration of AI agents into Slack with its new NanoClaw Slack integration, enabling users to create persistent teams of AI colleagues from a single message. Unlike previous attempts at AI integration that often felt clunky, NanoClaw allows for the effortless creation of specialized agents, each with custom skills, workflows, and even avatars.

Serval’s super agent Catalyst creates roving background agents to identify and fix IT issues before they’re ticketed
Serval is making its AI agent, Catalyst, generally available Thursday, empowering teams to automate enterprise workflows with unprecedented ease. This "super agent" analyzes ticket history, SOPs, and instructions to draft workflows, skills, and dashboards – even proactively identifying and fixing IT issues before they reach a ticket queue. Unlike competitors, Catalyst operates as a single administrative layer, moving from opportunity discovery to deploying proactive agents.

TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion than Claude Managed Agents
TrueFoundry introduces TrueForge, a new open-source AI agent harness designed to empower enterprise developers and reduce costs. Built by former Meta and Google engineers, TrueForge offers a vendor-neutral solution, compatible with various AI models and deployable across different infrastructures. Initial testing reveals impressive cost savings—up to 75% less than Anthropic’s Claude Managed Agents—achieved through intelligent context engineering.

Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation history locally
Block, the technology company behind Square and Cash App, is open-sourcing Berd, a desktop application designed to streamline AI agent workflows. Available now on GitHub under the Apache 2.0 license, Berd provides a unified workspace for users to manage projects, models, and conversation history locally. This innovative tool empowers users to explore AI capabilities across various agents and harnesses, offering a future-focused alternative to fragmented experiences.

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek’s V4 Flash, initially lauded as a "total monster" for its impressive leaderboard performance and remarkably low pricing, is experiencing a shift in perception. Recent testing reveals it completes only 53.8% of complex agent tasks in real-world scenarios. Simultaneously, DeepSeek is adjusting its pricing model, increasing rates by as much as 1,100% for certain token types.

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Google is accelerating AI innovation with the release of Gemini 3.7 Flash, its "most intelligent workhorse model yet" for coding and agentic workflows. This upgrade prioritizes diligent planning and disciplined execution, showing significant gains in debugging, web development, and enterprise automation—potentially reducing human intervention. Notably, Google is offering a 50% introductory price cut through the end of 2026, making it a compelling option for high-volume applications.

Why Capital One built its multi-agent AI platform around open-weight models
At VB Transform 2026, Capital One’s Kel Vanee detailed the bank’s strategic shift toward building AI, not just using it. Capital One constructed a scalable, multi-agent AI platform centered around deeply customized open-weight models, leveraging proprietary data for enhanced accuracy and extensibility. This approach, underpinned by prior investments in data transformation and cloud adoption, enables the bank to optimize workflows, from fraud detection to customer service, and even automate internal infrastructure tuning.

Constraining Output Space for SLM Narrow Automation Optimization
Optimizing narrow automation for Semantic Layer Models (SLMs) unlocks significant productivity gains. This series begins by exploring a crucial technique: constraining the output space, rather than solely relying on parsing generated text. By limiting potential outputs, we achieve greater efficiency and reliability in automated workflows. This initial article will detail how to implement this approach effectively. For broader context on navigating the evolving AI landscape, see our article, "New EU Guidelines For AI Labelling," for essential insights into regulatory considerations.

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges
Writer today unveiled Palmyra X6, a new AI agent model poised to significantly reduce costs for enterprise users. Paired with its rebuilt agent orchestration “harness,” Palmyra X6 delivers an average of 52% lower operational costs, alongside a 48% speed improvement and 10% quality boost. Leveraging a post-trained version of GLM-5.2, Writer emphasizes control and cost transparency, offering governance tools and multi-model support—a strategy echoing the shift towards pragmatic AI adoption, as explored in "Why Capital One built its multi-agent AI platform around open-weight models."

Skan AI raises $63 million betting that watching how employees actually work is the missing layer of enterprise AI
Skan AI has secured $63 million in Series C funding, co-led by Cathay Innovation and Dell Technologies Capital, signaling a significant bet on understanding how employees *actually* work. The company's approach diverges from traditional enterprise AI, which often falters due to a disconnect between documented processes and real-world execution. Skan builds a "context graph of work" by observing employee activity across applications, ultimately aiming to automate workflows and unlock substantial productivity gains—a strategy that echoes the foundational role CRM played in customer data management.

SpaceXAI's Grok Bot turns agents into persistent digital coworkers that can operate your apps for $120-per-month
SpaceXAI’s Grok Bot introduces a transformative approach to AI assistance, moving beyond simple prompts to continuously execute work within your existing applications—essentially creating persistent digital coworkers. Starting at $120 per month, this early beta version allows users to delegate tasks and workflows to Bots, which operate independently and can even hand off work to one another. Like OpenAI's recent focus on longer, multi-step tasks, Grok Bot aims to bridge the gap between near-completion and finished work, offering a new model for productivity.

Your AI agent may be ready. Your sales motion probably isn’t.
Your AI agent may be ready. Your sales motion probably isn’t. The shift to agent-guided buying is accelerating, with Gartner predicting 90% of B2B purchases will leverage AI by 2028. While many companies are investing in agent development, a critical gap remains: the time it takes to convert buyer interest into active customers. Companies winning now prioritize streamlined commerce, shrinking deal cycles from weeks to hours. Salesforce's AgentExchange addresses this friction, connecting discovery, commerce, and activation to empower faster, more efficient growth.

No cloud, no GPUs, no problem: Liquid AI's new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi
Liquid AI has unveiled LFM2.5-2.6B, a new open-weight language model designed to bring powerful AI agents to devices as small as a Raspberry Pi – a significant step toward accessible edge AI. This model, boasting 2.6 billion parameters and a 128,000-token context window, runs entirely on local hardware without cloud inference or GPUs, ideal for high-volume tasks like automation and connectivity-limited environments. Explore how this innovative solution transforms data management and expands possibilities for enterprises, as highlighted in our recent coverage of Qwen 3.8-Max.

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
Recent benchmarks of Qwen 3.8-Max and Claude Opus 5 highlight a crucial shift in evaluating large language models: raw benchmark scores don't accurately predict real-world costs. While initial marketing suggested Qwen 3.8-Max rivaled Claude, independent testing revealed significant performance variations tied to differing time budgets. The key takeaway? Adopt a "cost per successful task" metric, factoring in all attempts – including failures – to truly understand model efficiency.