AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost
Our take

To quote an ancient Jedi Master "Begun, the AI price wars have!" OpenAI’s dramatic price reductions for its GPT-5.6 models, particularly Luna’s 80% cut, signal a significant shift in the landscape of AI model deployment. This move, coupled with recent releases from Google and Anthropic, underscores a growing realization that access to powerful models alone isn’t sufficient; the economics of running those models in production are now paramount. It’s a fascinating development, especially considering that AI hedge fund Situational Awareness may have sold its public portfolio, but still holds onto its Anthropic shares, highlighting the continued, albeit shifting, valuations within the sector. This competition isn't just about bragging rights; it’s about making AI practical and scalable for businesses across a wide range of applications.
The price cuts themselves reveal a nuanced strategy. OpenAI isn't simply trying to undercut every competitor on a pure token basis – Xiaomi’s MiMo models, for instance, still offer lower per-token costs. Instead, they're strategically repositioning their GPT-5.6 series. Luna’s dramatic price drop targets the high-volume, cost-sensitive applications – summarization, classification, routing – where small cost differences can drastically impact overall expenses. Terra's modest reduction brings it closer to competitive offerings, and the introduction of Sol Fast mode caters to developers willing to pay a premium for reduced latency in demanding workloads. This tiered approach acknowledges that different use cases have different priorities, and OpenAI is attempting to capture market share across the spectrum. The fact that Mastercard spent decades training its fraud system to see bots as thieves illustrates the complex, often adversarial, relationship between businesses and AI, and underscores the need for efficient and cost effective AI solutions.
What’s particularly noteworthy is the contrasting approaches from OpenAI, Google, and Anthropic. OpenAI is directly lowering per-token rates, Google is focusing on optimizing token usage and tool calls to reduce overall cost, and Anthropic is emphasizing improved performance at a fixed price. All three strategies ultimately target the same goal: minimizing the total cost of completing production work. This signals a move beyond the initial hype around simply having access to the "best" model; the focus is now squarely on operational efficiency. The competition isn't about raw intelligence anymore, but about the ability to deliver useful results at a predictable and affordable price. The shift is a healthy one, and it should drive innovation in model architecture and deployment strategies.
Looking ahead, the AI price wars are likely to intensify. The trend of commoditization will continue, forcing providers to differentiate themselves not just on model performance, but also on factors like ease of integration, developer tooling, and support. The question now is whether these price cuts represent a sustainable business model for OpenAI, or a temporary measure to gain market share. More importantly, the increased affordability of frontier models will likely lead to a surge in experimentation and adoption across various industries. Will we see a wave of new AI-powered applications emerge, driven by the reduced cost of using these powerful tools? That's the development to watch—the democratization of AI intelligence and its impact on the world.
To quote an ancient Jedi Master "Begun, the AI price wars have!"
OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model.
The cuts place Luna much closer to the lowest-cost commercial models in the market and arrive just a few days after Anthropic released its highly performant Claude Opus 5 at the same price as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two rival models built around lower inference costs, faster execution and more efficient agent workloads.
OpenAI is successfully undercutting Google's price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost.
OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens.
Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14.
Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens.
The company says Fast mode delivers up to 2.5 times the throughput without changing the model’s underlying intelligence.
OpenAI co-founder and CEO Sam Altman took to X to announce the changes as "major price cuts today."
VentureBeat Frontier AI model API pricing comparison
Model | Input ($/1M) | Output ($/1M) | Total ($/1M) | Source |
MiMo-V2.5 Flash | $0.10 | $0.30 | $0.40 | |
deepseek-v4-flash | $0.14 | $0.28 | $0.42 | |
deepseek-v4-pro | $0.435 | $0.87 | $1.305 | |
GPT-5.6 Luna | $0.20 | $1.20 | $1.40 | |
MiniMax-M3 | $0.30 | $1.20 | $1.50 | |
LongCat-2.0 — limited-time promo | $0.30 | $1.20 | $1.50 | |
Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $1.75 | |
Qwen3.7-Plus | $0.40 | $1.60 | $2.00 | |
MiMo-V2.5 | $0.40 | $2.00 | $2.40 | |
Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $2.80 | |
LongCat-2.0 — standard | $0.75 | $2.95 | $3.70 | |
MiMo-V2.5 Pro (≤256K) | $1.00 | $3.00 | $4.00 | |
GLM-5.2 | $1.40 | $4.40 | $5.80 | |
Grok 4.5 | $2.00 | $6.00 | $8.00 | |
MiMo-V2.5 Pro (>256K) | $2.00 | $6.00 | $8.00 | |
Gemini 3.6 Flash | $1.50 | $7.50 | $9.00 | |
Qwen3.7-Max | $2.50 | $7.50 | $10.00 | |
Gemini 3.5 Flash | $1.50 | $9.00 | $10.50 | |
Gemini 3.1 Pro Preview (≤200K) | $2.00 | $12.00 | $14.00 | |
GPT-5.6 Terra | $2.00 | $12.00 | $14.00 | |
GPT-5.4 | $2.50 | $15.00 | $17.50 | |
Kimi K3 | $3.00 | $15.00 | $18.00 | |
Gemini 3.1 Pro Preview (>200K) | $4.00 | $18.00 | $22.00 | |
Claude Opus 5 | $5.00 | $25.00 | $30.00 | |
GPT-5.5 | $5.00 | $30.00 | $35.00 | |
GPT-5.5 Instant (chat-latest) | $5.00 | $30.00 | $35.00 | |
Sakana Fugu Ultra (≤272K) | $5.00 | $30.00 | $35.00 | |
GPT-5.6 Sol — Standard mode | $5.00 | $30.00 | $35.00 | |
Claude Fable 5 / Claude Mythos 5 | $10.00 | $50.00 | $60.00 | |
GPT-5.6 Sol — Fast mode | $10.00 | $60.00 | $70.00 |
Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers.
OpenAI moves Luna into the low-cost tier
The most consequential change is the Luna price cut.
When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40.
That places Luna below Google’s Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI’s own GPT-5.4 and Terra models by a wide margin.
It is not the cheapest model in the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market’s low-cost inference tier.
OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option.
The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost.
Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint.
Terra drops to match Google’s Gemini 3.1 Pro pricing
Terra’s 20% reduction moves its combined price from $17.50 to $14 per million tokens.
At that level, Terra now matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less.
It also undercuts OpenAI’s GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI's Nic Dunz noted on X:
The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard.
Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI’s decision to charge a premium for latency-sensitive workloads rather than lower Sol’s base price.
Cuts follow Google’s low-cost Gemini releases and Anthropic's Claude Opus 5
OpenAI’s pricing changes come only about a week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens.
Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks.
Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google’s 3.5 series.
However, OpenAI's models are more performant than Google's, according to third party analysis outfits like Artificial Analysis, with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI.
As AI coding startup Cognition noted on X, GPT-5.6 now "sits on the pareto curve of price/performance efficiency," posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market.
And yet, rival Anthropic's Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper.
The model costs $5 per million input tokens and $25 per million output tokens—the same rates as Opus 4.8—but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost.
Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings.
That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone.
The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI’s response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently.
The market shifts from model access to model economics
The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production.
OpenAI is still not the lowest-priced provider on a pure token basis. But Luna’s 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors.
That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows.
OpenAI’s latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company’s direct answer to the industry’s growing low-cost model segment.
Read on the original site
Open the publisher's page for the full experience