OpenAI cuts GPT-5.6 Luna and Terra API prices while adding faster Sol processing

OpenAI cuts GPT-5.6 Luna and Terra API prices while adding faster Sol processing

OpenAI lowered GPT-5.6 Luna and Terra API prices and replaced Priority Processing with Fast mode for Sol.

Format News Brief
Read Time 2 min
Category AI & Technology
Updated Aug 01, 2026

OpenAI has made a fresh price-performance move around its GPT-5.6 family, saying on July 30 that two of its lower-cost models are now cheaper for API customers and for usage counted inside ChatGPT Work and Codex. The company says GPT-5.6 Luna, described as its fastest and most affordable model in the family, now costs 80% less, while GPT-5.6 Terra, its balanced model for everyday work, costs 20% less.

The update matters because it is aimed less at headline benchmark records and more at the operating economics of AI in production. OpenAI is positioning Luna for high-volume work where businesses need tool use and multi-step workflows without paying frontier-model prices for every call. Terra remains the middle option for routine workplace and agent tasks where quality, latency, and cost all matter.

What changed

  • Starting July 30, Terra API pricing is $2 per million input tokens and $12 per million output tokens.
  • Luna API pricing is now $0.20 per million input tokens and $1.20 per million output tokens.
  • Sol pricing is unchanged, but a new Fast mode replaces Priority Processing in the API.
  • OpenAI says Fast mode gives GPT-5.6 Sol up to 2.5 times faster speeds than Standard processing at twice the price.

OpenAI also tied the customer-facing price changes to engineering work published a day earlier. In that post, the company said GPT-5.6 Sol helped optimize parts of its own inference stack, including production kernels, traffic routing, speculative decoding experiments, and workload-specific serving configurations. OpenAI says the kernel work reduced end-to-end serving costs by 20%, while improvements to token-generation efficiency added more than 15%.

For developers, the practical signal is that model choice is becoming more granular. A workflow may use a stronger model to plan or resolve uncertain work, then hand clearer implementation, testing, classification, or document-processing tasks to a cheaper model. OpenAI says existing API requests tagged for priority processing will continue to work and will use Fast mode instead, reducing migration friction for teams already using the premium latency path.

The claims are OpenAI's own and still need to be tested inside real customer workloads. But the direction is notable: the company is competing not only on model capability, but also on how cheaply and quickly those models can be served at scale.

Sources

Cover image: kewl, source, licensed under BY.

Comments (0)

Leave a Comment

Loading comments...