OpenAI has initiated a significant price reduction across its GPT-5.6 frontier series, slashing the cost of its smallest and fastest model, GPT-5.6 Luna, by a remarkable 80%, and its mid-tier model, GPT-5.6 Terra, by 20%. Simultaneously, the company has introduced a premium "Fast mode" for its flagship GPT-5.6 Sol model, signaling a strategic shift to aggressively compete on cost while enhancing performance options. These aggressive price cuts, particularly for Luna, bring OpenAI’s offerings into direct contention with the lowest-cost commercial AI models available in the market. This move arrives mere days after rivals Anthropic launched its highly performant Claude Opus 5 at the same price point as its predecessor, Opus 4.8, and Google introduced its Gemini 3.6 Flash and Gemini 3.5 Flash-Lite models, all engineered for lower inference costs, faster execution, and more efficient agent workloads. OpenAI is now not only aiming to undercut Google on a price-per-intelligence metric but also strategically attempting to entice Anthropic users by offering a compelling speed boost at a more attractive price.
OpenAI has detailed the new pricing structure, with Luna now priced at a highly competitive $0.20 per million input tokens and $1.20 per million output tokens, resulting in a combined input-plus-output cost of $1.40 per million tokens. The Terra model will see its combined cost reduced to $14 per million tokens, with input tokens priced at $2 million and output tokens at $12 million. The pricing for the Sol Standard model remains unchanged at $5 per million input tokens and $30 per million output tokens. To cater to latency-sensitive applications, OpenAI is introducing the Sol Fast mode, which will be priced at double the Standard rate: $10 per million input tokens and $60 per million output tokens. The company asserts that this Sol Fast mode can deliver up to 2.5 times the throughput without compromising the model’s underlying intelligence, effectively offering a significant speed upgrade for a premium. OpenAI co-founder and CEO Sam Altman announced these changes via X (formerly Twitter), labeling them as "major price cuts today," underscoring the company’s aggressive stance in the evolving AI landscape.
| Model | Input ($/1M) | Output ($/1M) | Total ($/1M) | Source |
|---|---|---|---|---|
| MiMo-V2.5 Flash | $0.10 | $0.30 | $0.40 | Xiaomi |
| deepseek-v4-flash | $0.14 | $0.28 | $0.42 | DeepSeek |
| deepseek-v4-pro | $0.435 | $0.87 | $1.305 | DeepSeek |
| GPT-5.6 Luna | $0.20 | $1.20 | $1.40 | OpenAI |
| MiniMax-M3 | $0.30 | $1.20 | $1.50 | MiniMax |
| LongCat-2.0 – limited-time promo | $0.30 | $1.20 | $1.50 | LongCat |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $1.75 | |
| Qwen3.7-Plus | $0.40 | $1.60 | $2.00 | Alibaba Cloud |
| MiMo-V2.5 | $0.40 | $2.00 | $2.40 | Xiaomi |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $2.80 | |
| LongCat-2.0 – standard | $0.75 | $2.95 | $3.70 | LongCat |
| MiMo-V2.5 Pro (≤256K) | $1.00 | $3.00 | $4.00 | Xiaomi |
| GLM-5.2 | $1.40 | $4.40 | $5.80 | Z.ai |
| Grok 4.5 | $2.00 | $6.00 | $8.00 | xAI |
| MiMo-V2.5 Pro (>256K) | $2.00 | $6.00 | $8.00 | Xiaomi |
| Gemini 3.6 Flash | $1.50 | $7.50 | $9.00 | |
| Qwen3.7-Max | $2.50 | $7.50 | $10.00 | Alibaba Cloud |
| Gemini 3.5 Flash | $1.50 | $9.00 | $10.50 | |
| Gemini 3.1 Pro Preview (≤200K) | $2.00 | $12.00 | $14.00 | |
| GPT-5.6 Terra | $2.00 | $12.00 | $14.00 | OpenAI |
| GPT-5.4 | $2.50 | $15.00 | $17.50 | OpenAI |
| Kimi K3 | $3.00 | $15.00 | $18.00 | Moonshot AI |
| Gemini 3.1 Pro Preview (>200K) | $4.00 | $18.00 | $22.00 | |
| Claude Opus 5 | $5.00 | $25.00 | $30.00 | Anthropic |
| GPT-5.5 | $5.00 | $30.00 | $35.00 | OpenAI |
| GPT-5.5 Instant (chat-latest) | $5.00 | $30.00 | $35.00 | OpenAI |
| Sakana Fugu Ultra (≤272K) | $5.00 | $30.00 | $35.00 | Sakana AI |
| GPT-5.6 Sol – Standard mode | $5.00 | $30.00 | $35.00 | OpenAI |
| Claude Fable 5 / Claude Mythos 5 | $10.00 | $50.00 | $60.00 | Anthropic |
| GPT-5.6 Sol – Fast mode | $10.00 | $60.00 | $70.00 | OpenAI |
Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers.
The most impactful aspect of OpenAI’s announcement is the dramatic price cut for GPT-5.6 Luna. Previously, Luna was priced at $1 per million input tokens and $6 per million output tokens, totaling $7 million per million tokens. The new pricing has drastically reduced this combined figure to $1.40 per million tokens. This strategic repositioning places Luna significantly below Google’s Gemini 3.5 Flash-Lite, which has a combined cost of $2.80 per million tokens, and substantially undercuts Gemini 3.6 Flash at $9 per million tokens. Furthermore, Luna now boasts a lower price point than OpenAI’s own older GPT-5.4 and Terra models. While Luna isn’t the absolute cheapest model on the market – with providers like Xiaomi (MiMo-V2.5 Flash) and DeepSeek (flash model) offering lower per-token rates – this reduction effectively catapults an OpenAI frontier-model into direct competition with the industry’s most cost-effective inference solutions. OpenAI categorizes the GPT-5.6 series as its frontier model family, with Sol representing the pinnacle of capability, Terra serving as the balanced mid-tier option, and Luna positioned as the smallest and fastest model. The GPT-5.6 lineup was initially unveiled in late June 2026, following a restricted rollout at the request of the U.S. government, before wider availability. Each model within the series is designed to offer distinct trade-offs between intelligence, latency, and cost. Sol is optimized for the most demanding reasoning-intensive and agentic workloads, including complex coding, multi-step planning, and sophisticated tool-using systems. Terra is engineered for general production environments where a blend of performance and efficiency is paramount. Luna, on the other hand, is tailored for high-throughput, low-latency tasks such as summarization, classification, intelligent routing, and the development of lightweight real-time assistants, where minimizing cost per request is the primary consideration.

The 20% price reduction for Terra brings its combined input-and-output cost down from $17.50 per million tokens to $14 per million tokens. This new pricing benchmark for Terra now aligns it with Google’s Gemini 3.1 Pro Preview pricing for contexts up to 200,000 tokens. It also significantly undercuts OpenAI’s own GPT-5.4 model, which remains priced at $2.50 per million input tokens and $15 per million output tokens. As Krea AI’s Nic Dunz highlighted on X, this means users can now access comparable intelligence from Terra at approximately 1/13th of the cost of GPT-5.4. This adjustment effectively widens the price differentiation between OpenAI’s three GPT-5.6 tiers. Luna now costs one-tenth of Terra’s combined price, while Terra is 60% less expensive than Sol Standard. Conversely, Sol Fast mode moves in the opposite direction, with its combined cost of $70 per million tokens positioning it as the most expensive model configuration within the provided comparison. This pricing strategy reflects OpenAI’s deliberate choice to command a premium for enhanced latency performance, rather than reducing the base price of the Sol model.
These substantial pricing adjustments by OpenAI occur within a rapidly evolving competitive landscape, closely following Google’s recent introduction of its own cost-effective Gemini 3.6 Flash and Gemini 3.5 Flash-Lite models. Google had priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens, while Gemini 3.5 Flash-Lite was set at $0.30 per million input tokens and $2.50 per million output tokens. Google positioned both models as being optimized for agent deployment economics, emphasizing that reduced token usage, fewer reasoning steps, and decreased tool calls could collectively lower the total cost of long-running tasks in software engineering and knowledge work. Reports suggest that Gemini 3.6 Flash utilizes 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% for certain extended engineering workloads. Gemini 3.5 Flash-Lite is marketed as the fastest model within Google’s 3.5 series. However, third-party analyses from entities like Artificial Analysis indicate that OpenAI’s GPT-5.6 models exhibit superior performance compared to Google’s Gemini offerings. Even the more affordable Luna model reportedly outperforms Gemini 3.6 Flash and the older Gemini 3.1 Pro model, presenting a significantly more favorable cost-per-intelligence ratio for OpenAI. As noted by AI coding startup Cognition on X, GPT-5.6 now "sits on the pareto curve of price/performance efficiency," illustrating that these models offer among the highest intelligence for the lowest cost in the market.
Despite these competitive advancements, Anthropic’s Claude Opus 5 remains a formidable competitor, offering performance comparable to GPT-5.6 Sol but at a slightly lower price point. Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens – the same rates as its predecessor, Opus 4.8. Anthropic claims that Opus 5 delivers nearly the full intelligence of its more advanced Fable 5 model at roughly half the cost. Unlike OpenAI’s direct price cuts for Luna and Terra, Anthropic has maintained the Opus API sticker price. Instead, they have effectively lowered the price per unit of capability by replacing Opus 4.8 with a more advanced model at the same $30 combined input-and-output rate. Anthropic has also introduced an adjustable effort setting, allowing developers to fine-tune the balance between reasoning depth, speed, and token consumption. This nuanced approach to pricing and performance optimization is critical for enterprise buyers. OpenAI’s strategy focuses on direct per-token rate reductions, Google emphasizes cost savings through reduced token usage and tool calls, and Anthropic highlights enhanced task performance at a consistent price. All three approaches ultimately target the same operational metric: the total cost of completing production work, rather than solely the advertised cost of an individual token. The timing of these moves underscores the rapid evolution of pricing as a key competitive differentiator among leading AI model providers. OpenAI’s latest adjustments are not tied to a new model generation but rather focus on optimizing the economics of deploying its recently released GPT-5.6 series.
The recent pricing shifts signal a market transition where access to frontier-level AI capabilities is no longer the sole competitive battleground. The crucial question for enterprises now revolves around the cost-effectiveness and predictability of running these models in production environments. While OpenAI may not be the absolute lowest-priced provider on a per-token basis, the substantial 80% reduction in Luna’s price fundamentally alters its market position. Luna has transitioned from a mid-tier offering to a model directly competing within the pricing tier occupied by smaller models from Google, Xiaomi, DeepSeek, MiniMax, and other vendors. This is particularly significant for high-volume applications, where even minor differences in token pricing can lead to substantial cost savings across a wide array of use cases, including coding agents, document analysis systems, internal search tools, and automated workflows. Therefore, OpenAI’s latest strategic move appears less like a routine pricing adjustment and more like a deliberate repositioning of its GPT-5.6 series. Sol continues to serve as the premium, high-performance option. Terra is now more competitively positioned against other pro-tier systems. And Luna is effectively OpenAI’s direct response to the burgeoning segment of low-cost, high-efficiency AI models in the industry. This competitive dynamic is likely to drive further innovation and price adjustments as companies vie for market share in the rapidly expanding AI ecosystem.

