LLMs

Shortening prompts costs more; asking for brevity saves.

Recent research definitively answers a critical question: does instructing an LLM to "be concise" actually save money?

3 min readMachine Learning

The results are in, and they're not what most teams expect. A new study tested nine models across five reduction levels, measuring what happens when you compress the input prompt versus when you instruct the model to write a shorter output. The finding is refreshingly direct: shortening the output saves money and holds accuracy steady, often cutting costs by 1.5x on average and up to 3x in the best case. Shortening the input prompt does the opposite. On the worst benchmark, it cost 96% more because the model simply answers longer to fill the gap you created, and accuracy drops. You pay more and get worse answers. That is a counterintuitive result worth sitting with.

For anyone building on top of these APIs, this is the kind of insight that changes how you design prompts. We have covered the practical side of Unlock LLM Training: A Practical Guide to Distributed Algorithms and the decision-making trade-offs in Jev vs LLMs: Evaluating AI for Practical Decision-Making, but this study gets at something more immediate: the cost structure itself. Output tokens are more expensive than input tokens, so prompting for fewer output tokens is a direct lever on your bill. The data confirms it works across languages too, from English to Swahili to Thai, which suggests this is not a quirk of one tokenizer or one model family. It is a general property of how these systems behave.

The nuance, however, is worth flagging. When the shortened output is correct, about half the time it no longer matches how the model would have reasoned without the constraint. That is fine if you only care about the final answer, and many use cases do. But if you are relying on the model's reasoning trace for auditability, for debugging, or for teaching, you are losing more than you might realize. The cost savings come with a hidden trade-off in explainability. And with providers starting to ship their own "concise" options, the black box gets deeper. We cannot see how they price those features, so we do not actually know if the savings reach your account or stay with the vendor. The study is honest about that limitation.

Here is the concrete takeaway we would give anyone asking: if you control the prompting yourself via the API, you can safely push for shorter outputs and pocket the savings without sacrificing accuracy. But do not touch the input prompt to cut costs, because the model will compensate by spending more on its own reply. The open question to watch is whether the managed "concise" modes from providers will pass the savings through to users or quietly keep the margin. That is the detail that will determine whether this becomes a standard cost-optimization practice or just another vendor feature. For now, the manual approach works, and the data is on your side.

From Machine Learning

LLMs are too verbose and with a black box model the only things you control are what goes in and how you tell it to write back. Yesterday Claude Code shipped a "concise output style" where Claude keeps things short. We already have a paper out about this!

We tested both channels, shortening the input prompt versus telling the model to output answer shorter, on the same questions across five reduction levels, and scored cost, accuracy, and whether the shortened text still matched what the model would have said unconstrained.

Read the original at Machine Learning