DeepSeek

DeepSeek on Beyond Market Intelligence: a running collection of 6 stories we have gathered and hand-picked because they are worth your time. Every post here touches on deepseek in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around deepseek, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

GLM-5.3-Flash will likely handle 45% of your AI workloads
VentureBeat

GLM-5.3-Flash will likely handle 45% of your AI workloads

GLM-5.3-Flash is poised to reshape AI workflows, potentially handling as much as 45% of your organization's workloads. This surprisingly capable model, recently revealed to be from Z.ai and running on Chinese infrastructure, delivers exceptional performance at a significantly lower cost – approximately nine cents per task compared to 67 cents for a comparable US mid-tier like GPT-5.6 Sol. With open weights and accessible inference options, GLM-5.3-Flash presents a compelling opportunity to optimize AI spending and accelerate development, as highlighted by Uber's recent cost-cutting measures.

Machine Learning

Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]

Recent research definitively answers a critical question: does instructing an LLM to "be concise" actually save money? Across nine models—including GPT-4o and Claude Haiku—our analysis reveals a clear winner: prompting for shorter output consistently reduces costs by 1.5x on average (up to 3x in some cases) while maintaining accuracy. Conversely, shortening input prompts proved counterproductive, increasing costs and diminishing answer quality. This highlights a key insight: controlling output tokens is the most effective strategy for cost optimization, as demonstrated in our paper.

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
VentureBeat

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

DeepSeek’s V4 Flash, initially lauded as a "total monster" for its impressive leaderboard performance and remarkably low pricing, is experiencing a shift in perception. Recent testing reveals it completes only 53.8% of complex agent tasks in real-world scenarios. Simultaneously, DeepSeek is adjusting its pricing model, increasing rates by as much as 1,100% for certain token types.

DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices
VentureBeat

DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

DeepSeek is expanding beyond model development, launching DeepSeek Harness v0.1, an open-source agent harness designed as an alternative to tools like Anthropic’s Claude Code. Alongside this, the company released DeepSeek-V4-Pro, an updated flagship model optimized for agentic workloads, now accessible via DeepSeek’s web interface, mobile app, and API. While V4-Pro offers enhanced capabilities and OpenAI Responses API support, developers should note a shift to peak and off-peak API pricing, beginning Sunday, Aug. 16, which will substantially impact costs.

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size
VentureBeat

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size

Poolside's Laguna S 2.1 introduces a compelling new option in the open-weight coding model landscape. This 118-billion-parameter system, activating just 8 billion parameters per token, impressively matches or surpasses models many times its size on agentic coding tasks, achieving top scores on benchmarks like Terminal-Bench 2.1. With a permissive OpenMDW-1.1 license and broad ecosystem support, Laguna S 2.1 represents a strategic move to empower Western users with trustworthy, self-hostable AI.

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems
VentureBeat

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems

Moonshot AI has unveiled Kimi K3, a 2.8-trillion-parameter model now recognized as the world’s largest open-source AI, rivaling top proprietary systems from Anthropic and OpenAI. This release, timed before the 2026 World Artificial Intelligence Conference, marks a significant moment in the global AI race and a remarkable comeback for the Beijing-based startup. Full model weights will be released July 27th, allowing users to explore its capabilities—and potentially reshape their data strategies—at kimi.com.