GPT-5.4
GPT-5.4 on Beyond Market Intelligence: a running collection of 5 stories we have gathered and hand-picked because they are worth your time. Every post here touches on gpt-5.4 in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around gpt-5.4, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.
Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]
Recent research definitively answers a critical question: does instructing an LLM to "be concise" actually save money? Across nine models—including GPT-4o and Claude Haiku—our analysis reveals a clear winner: prompting for shorter output consistently reduces costs by 1.5x on average (up to 3x in some cases) while maintaining accuracy. Conversely, shortening input prompts proved counterproductive, increasing costs and diminishing answer quality. This highlights a key insight: controlling output tokens is the most effective strategy for cost optimization, as demonstrated in our paper.

Webwright: Why AI Web Agents Should Write Code, Not Click
For years, web agents have struggled with complex, long-horizon tasks, relying on a sequential click-by-click approach. Microsoft Research’s Webwright offers a transformative alternative: empowering AI models to write code directly. This shift, granting the model a terminal, yields impressive results, boosting success rates from 33.5% to 60.1% on challenging tasks. Unlike traditional agents that leave behind only a click trace, Webwright produces reusable command-line tools.

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff
Hark, a new AI startup founded by serial entrepreneur Brett Adcock, introduces Handoff, a computer use agent (CUA) poised to transform how we interact with the open web. Achieving a leading 97.7 score on the Online-Mind2Web benchmark—outperforming models like GPT-5.4 and Claude Opus 4.8—Handoff offers autonomous task completion, from online ordering to candidate outreach. With significantly lower operational costs, Hark empowers users to explore a future where AI handles routine digital tasks. Sign-ups are open now at hark.
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A recent solo evaluation project rigorously assessed six frontier LLMs—GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, and Grok 4.3—across eight established bias benchmarks, encompassing over 20,600 examples. Findings reveal a consistent leftward political leaning among all models except Grok, despite its self-reported right-leaning stance. Notably, GPT-5.4 exhibited the highest refusal rate (20.3%) when addressing race-related inquiries requiring explicit racial identification. For deeper insights into AI memory systems, explore "Context Windows Forget What Matters." Full data and

Microsoft launches AI cybersecurity model, agentic defense platform to cut enterprise security costs
Microsoft is reshaping enterprise cybersecurity with the launch of MAI-Cyber-1-Flash, a compact AI model embedded within the agentic defense platform, MDASH. This innovative system, achieving 96% accuracy on the CyberGym benchmark, delivers significant cost savings—roughly 50%—compared to existing configurations. Project Perception, a coordinating agentic security system, enters public preview August 3rd. Microsoft’s approach prioritizes cost-effective solutions, leveraging a specialized model for routine tasks and OpenAI's GPT-5.