Input tokens
Input tokens on Beyond Market Intelligence: a running collection of 3 stories we have gathered and hand-picked because they are worth your time. Every post here touches on input tokens in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around input tokens, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.
Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]
Recent research definitively answers a critical question: does instructing an LLM to "be concise" actually save money? Across nine models—including GPT-4o and Claude Haiku—our analysis reveals a clear winner: prompting for shorter output consistently reduces costs by 1.5x on average (up to 3x in some cases) while maintaining accuracy. Conversely, shortening input prompts proved counterproductive, increasing costs and diminishing answer quality. This highlights a key insight: controlling output tokens is the most effective strategy for cost optimization, as demonstrated in our paper.

GLM-5.3 hits the API at $1.4/$4.4 per million tokens
Z.ai has made GLM-5.3, its new open-source language model boasting advanced coding and agent capabilities, accessible via API. Developers can now integrate this frontier model into their applications at a competitive rate of $1.40 per million input tokens and $4.40 per million output tokens—unchanged from its predecessor, GLM-5.2. Independent benchmarks place GLM-5.3 among the world’s top open-weight models, demonstrating strong performance at a notably lower cost than premium alternatives. For teams exploring coding and agent workloads, GLM-5.3 represents a compelling, accessible option.

AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost
The AI landscape is rapidly evolving, and the latest development is a full-blown price war. OpenAI has sharply reduced prices on its GPT-5.6 models, cutting Luna by a striking 80% and Terra by 20%, effectively undercutting competitors like Google and Anthropic. This strategic move, announced by Sam Altman, positions Luna competitively within the low-cost inference tier and underscores a shift toward model economics as the key differentiator.