OpenAI Cuts GPT-5.6 Luna Pricing by 80%
OpenAI announced the pricing update on July 30, following improvements to its models, inference infrastructure and the systems used to manage agent-based workflows. According to the company, these changes have reduced the amount of time, computing power and tokens required to complete many tasks.
Under the new standard API pricing for short-context requests, GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra is priced at $2 per million input tokens and $12 per million output tokens.
GPT-5.6 Sol pricing remains unchanged at $5 per million input tokens and $30 per million output tokens. Long-context requests are priced separately and remain more expensive than the standard rates.
The reductions are also reflected in ChatGPT Work and Codex usage. OpenAI has not reduced subscription prices or changed the overall quota budgets for those products. Instead, requests using Terra and Luna now consume fewer credits, allowing users to complete more work within existing plans.
Fast Mode Makes GPT-5.6 Sol Up to 2.5 Times Faster
OpenAI also introduced Fast mode for GPT-5.6 Sol in the API. It replaces the company’s previous Priority Processing option and is designed for workloads where response time is more important than the lowest possible cost.
Fast mode can run GPT-5.6 Sol at up to 2.5 times the speed of Standard processing. It costs twice the standard rate but uses the same underlying model, meaning OpenAI says there is no reduction in intelligence or reasoning capability.
The change is backward compatible. API requests already configured with the priority service tier will automatically use Fast mode, while developers can also select the new fast service tier directly.
This gives companies a clearer choice between lower-cost standard processing and faster execution for latency-sensitive applications such as interactive coding, customer support, financial research and time-critical agent workflows.
GPT-5.6 Models Target Different Workloads
The GPT-5.6 family became generally available earlier in July and includes three models aimed at different levels of capability, speed and cost.
- GPT-5.6 Sol is the flagship model for complex professional work, including advanced coding, cybersecurity, scientific research, computer use and long-running agent tasks.
- GPT-5.6 Terra is positioned as the balanced option for everyday business workloads where companies need strong reasoning without paying Sol-level prices.
- GPT-5.6 Luna is the fastest and least expensive model in the family, designed for cost-sensitive and high-volume workloads.
AI Wire Media previously examined the broader model family in OpenAI Launches GPT-5.6 and ChatGPT Work as AI Competition Intensifies.
Luna Expands the Range of Economically Viable AI Tasks
The biggest change is the new price of Luna. At $0.20 per million input tokens, the model can be used for tasks that would previously have been too expensive to run continuously or across millions of records.
Potential applications include document analysis, customer-interaction classification, routine software implementation, test generation, data extraction and multi-step background automation.
OpenAI says Luna can use tools and complete agent-based workflows rather than functioning only as a basic text-generation model. This makes it possible to use the lower-cost model for parts of a larger process while reserving Sol for decisions that require more advanced reasoning.
For example, Sol could define the architecture of a software project or resolve an ambiguous technical problem. Luna could then implement clearly specified changes, write tests, process files and review the results at a much lower cost.
OpenAI Moves Auto-Review to GPT-5.6 Luna
OpenAI also said it is upgrading the Auto-review feature in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna. Combined with Luna’s new pricing, the company expects the feature to cost approximately ten times less to operate.
Auto-review examines approval requests and proposed actions during agent-based coding workflows. Lowering the cost of this process could make continuous review more practical for software teams running large numbers of automated tasks.
The update fits OpenAI’s broader strategy of using smaller models for routine verification and implementation while directing difficult planning and reasoning work to more capable models.
OpenAI Claims Major Price-Performance Gains
According to OpenAI’s estimates, GPT-5.6 Luna provides performance comparable to models that were considered frontier-class about a year ago while costing roughly six cents for every dollar previously spent per task. The company also says Luna completes comparable work at nearly nine times the speed.
On Agents’ Last Exam, a benchmark focused on long-running professional workflows, OpenAI reports that Luna outperformed Claude Fable 5 while carrying an estimated cost per task that was nearly 99% lower.
These figures should be treated as company-reported benchmark estimates rather than a guarantee of equivalent savings in every production environment. Actual costs will depend on prompt size, reasoning settings, output length, tool use, context size and the number of retries required.
Prompt design can also have a significant impact. AI Wire Media previously covered OpenAI’s guidance on reducing unnecessary instructions in OpenAI Says Shorter GPT-5.6 Sol Prompts Cut Tokens and Costs.
What the Price Cuts Mean for Enterprise AI
The update reflects a wider shift in the AI market. Competition is no longer focused only on which company has the highest benchmark score. Model providers are increasingly competing on the total cost of completing useful work, including latency, token consumption, reliability and the number of steps required.
Lower-cost models such as Luna could allow companies to apply AI to a much larger number of routine operations. Terra offers a middle ground for tasks that need stronger reasoning, while Sol remains available for the most complex workflows and can now be accelerated through Fast mode when time matters.
For businesses, the main challenge will be selecting the right model for each stage of a workflow rather than sending every task to the most capable and expensive option. OpenAI is effectively encouraging customers to build systems that move between Sol, Terra and Luna according to the difficulty and value of the work.
ES
EN