tool calls
tool calls on Beyond Market Intelligence: a running collection of 7 stories we have gathered and hand-picked because they are worth your time. Every post here touches on tool calls in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around tool calls, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Software engineers' new job isn't writing code — it's designing the boundaries AI agents can't break
The role of the software engineer is evolving. As AI agents increasingly handle code generation—producing initial implementations of pipelines and integrations with remarkable speed—the focus shifts from syntax to system boundaries. Rather than crafting every line of logic, engineers are now tasked with designing robust frameworks where agent-generated code can thrive. This means establishing clear data contracts and feedback loops to ensure accuracy and prevent operational entropy, ultimately transforming the engineer’s value into the design of reliable, trustworthy systems.

Identity and permissions aren’t enough to govern AI agent behavior
Enterprise AI agent security demands a shift beyond traditional identity and permissions. While access controls remain foundational, they don't govern *how* an agent behaves once active, potentially turning legitimate access into unintended consequences at machine speed. Heather Ceylan, CISO at Box, emphasizes a layered approach that includes governing execution, ensuring permissions are dynamically scoped to the task at hand. Addressing this challenge requires a focus on content-level visibility, as highlighted in our recent article on Uber’s GitFarm, to secure the rapidly evolving AI landscape.

AWS Open-Sources Dogwood, Extending Cedar to Govern Sequences of Agent Tool Calls
AWS has expanded its policy governance capabilities with the open-source release of Dogwood, an extension of Cedar. Dogwood introduces temporal reasoning, enabling rules to evaluate sequences of agent tool calls—crucial for approvals, rate limits, and running totals. Released under Apache 2.0, Dogwood is initially supported within AgentCore Policy. While the reference interpreter isn’t production-ready, this marks a significant step towards more sophisticated agent control. For context on broader agent tracing implementations, see our coverage of Cloudflare’s recent agent tracing launch.

NVIDIA Nemotron 3.5 Lightning: The AI Agent Workhorse
AI agents face a critical efficiency challenge: routine execution consumes the majority of their time. While frontier reasoning models excel at complex tasks, repeatedly applying them to simple actions—hundreds of tool calls, file operations, and validations—becomes slow and costly. NVIDIA’s Nemotron 3.5 Lightning addresses this directly, optimizing agent performance by intelligently allocating resources. Discover how this innovation transforms AI agent workflows, ensuring powerful reasoning is reserved for where it’s truly needed. For further insights into on-device agentic models, explore our article on Meta's Muse Glimmer.

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Google is accelerating AI innovation with the release of Gemini 3.7 Flash, its "most intelligent workhorse model yet" for coding and agentic workflows. This upgrade prioritizes diligent planning and disciplined execution, showing significant gains in debugging, web development, and enterprise automation—potentially reducing human intervention. Notably, Google is offering a 50% introductory price cut through the end of 2026, making it a compelling option for high-volume applications.
![I created an autonomous boxing benchmark [D]](https://preview.redd.it/r2i8f52ub8hh1.jpg?width=140&height=78&auto=webp&s=5ea73e9fad702339bb34f2c4c3a5ff60f2b2653b)
I created an autonomous boxing benchmark [D]
Introducing a novel AI benchmark: autonomous boxing. We've created a dynamic, physics-based environment where LLMs engage in simulated street fights, testing decision speed, adaptability, and strategic thinking. Models, like those utilizing Gemini-Flash-Live, can even dodge and counter punches. Currently tracking metrics like latency, action quality, and contextual awareness, we're seeking input on additional valuable stats to enhance this fun and insightful evaluation tool. For a deeper exploration of LLM training techniques, see our recent article, "Deep Dive on RL and OPD for Training LLMs."

Gemini 3.6 Flash Is Here: The Efficiency Release
While the industry awaited Gemini 3.5 Pro, Google quietly released Gemini 3.6 Flash on July 21, 2026—an efficiency-focused update to its speed tier. This release prioritizes streamlined performance, achieving comparable thinking capabilities to 3.5 Flash while reducing token usage, tool calls, and overall processing demands. It’s a practical step forward, demonstrating a commitment to optimized AI workflows. Explore the implications of this shift, and how it impacts agentic AI strategies—as discussed in our article, "Agentic AI vs AI Automation."