OpenAI has significantly reduced the prices of two of its latest artificial intelligence models, escalating competition in the generative AI market as leading developers battle to deliver more computing power at lower cost and convince businesses that AI investments can generate measurable returns.
The latest pricing overhaul is part of a growing industry shift that has seen the competitive edge being determined not only by model intelligence but also by the economics of deploying AI at scale. As enterprises expand AI adoption, they are placing greater emphasis on reducing inference costs, improving efficiency and lowering the total cost of ownership.
Chief Executive Officer Sam Altman announced the reductions on Thursday, saying OpenAI wants to offer “the best price/intelligence tradeoff at every level.”
The company slashed the cost of GPT-5.6 Luna by 80%, reducing pricing to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra received a 20% price cut, bringing costs down to $2 per million input tokens and $12 per million output tokens.
OpenAI also introduced a new “Fast” mode for GPT-5.6 Sol, its flagship frontier model. The option delivers up to 2.5 times faster performance through the API for twice the price while maintaining the same level of intelligence, giving developers the flexibility to prioritize speed for latency-sensitive applications.
The lower pricing extends beyond application programming interface (API) customers. OpenAI said Luna and Terra’s reduced costs will also be reflected in how usage is calculated for paid subscribers using Codex and ChatGPT Work, making advanced AI capabilities more affordable for software engineering, automation and enterprise workflows.
The company attributed the lower prices to improvements across its AI stack rather than reductions in model capability.
“Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context,” OpenAI said in a statement.
The company said advances in model architecture, optimized inference systems, improved routing, and better context management have enabled it to generate responses using fewer computing resources while maintaining performance. More efficient routing ensures hardware utilization remains high, while smarter context management allows AI agents to avoid repeating completed tasks, reducing unnecessary token generation and computing costs.
The announcement comes roughly three weeks after OpenAI launched the GPT-5.6 family following a temporary pause in the broader rollout at the request of the U.S. government. The models represent the company’s latest generation of AI systems designed for coding, reasoning and enterprise applications.
AI Pricing Emerges As The Industry’s Next Battleground
Over the past two years, model providers competed primarily by releasing increasingly capable AI systems. That contest is now evolving into a battle over who can deliver the highest intelligence at the lowest cost, a transition driven by rising enterprise AI adoption and surging demand for computing infrastructure.
Every AI query requires expensive graphics processing units (GPUs), electricity and networking resources. Lowering inference costs enables providers to improve margins while making their services more attractive to businesses that process billions of tokens each month.
Jacob Bourne, senior analyst at EMARKETER, said the latest move reflects growing pressure from enterprise customers seeking stronger returns on AI spending.
“The era of tokenmaxxing is over,” Bourne said.
“Enterprises have figured out how easy it is to burn tokens without getting value back, and they’re pushing back on those increasing AI bills.”
Organizations are now evaluating AI projects based on productivity gains and financial returns rather than raw model capability, forcing AI providers to optimize both performance and operating costs.
OpenAI’s pricing move is believed to have also been spurred by growing competition in the AI industry.
Chinese startup Moonshot AI recently released its open-weight Kimi K3 model, intensifying pressure on proprietary AI developers such as OpenAI by offering developers greater flexibility and lower deployment costs. Open-source and open-weight models continue to gain traction among enterprises seeking greater control over infrastructure and long-term operating expenses.
Anthropic, another leading AI developer, has been balancing subscription and usage-based pricing while managing finite computing capacity, illustrating the industry’s challenge of expanding access without overwhelming available infrastructure.
Google has likewise emphasized affordability as a competitive advantage, repeatedly highlighting the cost efficiency of its Gemini models as customers increasingly scrutinize AI spending.
Microsoft has adopted a similar strategy. During the company’s quarterly earnings call on Wednesday, Chief Executive Officer Satya Nadella repeatedly emphasized “cost efficiency” as a defining characteristic of Microsoft’s MAI-Thinking-1 model.
“We are building a new model system where the harness, context, memory, and action space are separate from any one model family, thereby moving the frontier on the cost to outcome curve,” Nadella said.
These point to a broader architectural shift across the AI industry. Rather than relying solely on larger models, developers are separating reasoning engines from memory, context management and external tools, allowing AI systems to complete more sophisticated tasks while using computing resources more efficiently.
Overall, OpenAI’s latest pricing cuts signal that the economics of artificial intelligence are becoming as important as the capabilities of the models themselves.
As AI applications move deeper into software development, customer service, research, finance and enterprise automation, businesses are demanding predictable operating costs alongside high performance. Lower token prices reduce barriers to adoption, encourage larger AI workloads and strengthen customer retention, particularly as organizations deploy AI agents that continuously process large volumes of data.






