Skip to content
← Back to Insights

GPT-6 Sol and Luna Cut Prices: Long-Task Savings Depend on What Gets Reused

AI OpenAI AI Agents Product Design News

TL;DR

OpenAI launches GPT-6 Sol and Luna with input prices 50% below prior promotional rates. Cached reads bring further discounts, but writes, retries and changing information still affect the total bill.

GPT-6 Sol and Luna Cut Prices: Long-Task Savings Depend on What Gets Reused

OpenAI has launched GPT-6 Sol and Luna with input prices 50% below their predecessors’ promotional rates. Public information still lacks comparisons of ordinary teams’ bills for complete tasks. If cheaper agents encourage additional rounds of work, lower unit prices will not necessarily reduce spending by the same proportion. Actual usage and accepted results must answer that question together. Prices

The event is dated 2026-09-22 in the United States. MacRumors published at 13:58 PDT that day, which was already 04:58 on September 23 in Taipei. This article carries the Taipei date of September 23 and a 12:43 reporting cutoff. This is coverage of a newly released model, not an attempt to equate the report’s timestamp with the moment of availability; OpenAI provides no precise publication time. Independent report

Standard input and output prices per million tokens are $2 and $10 for Sol, and $0.10 and $0.50 for Luna. The comparison is with promotional prices for the corresponding GPT-5.6 models. Both Sol rates are halved; Luna’s output rate falls further, from $1.20 to $0.50. Pricing table

Paid Plus, Pro, Business, Enterprise and Edu users can use both models in ChatGPT Work and Codex. Free and Go users can access Luna in the desktop application. Ordinary Chat mode is not yet supported, and the ChatGPT rollout is gradual. The API identifiers are gpt-6-sol and gpt-6-luna; a product announcement does not mean simultaneous availability across every interface. Access details

Cached reads save money, but writes cost extra

Long tasks repeatedly include the same background information. GPT-6 caching retains eligible matching prompt prefixes for at least 30 minutes after the latest write or reuse. Reads cost 90% less than standard input, while cache writes cost 1.25 times the standard input rate. The discount applies only to reused input, not to output or the entire task. Billing documentation

In OpenAI’s announcement, GitHub says the share of prompt tokens requiring fresh processing fell by more than 50% across billions of requests over recent months. That supplies evidence of real usage scale, but remains a partner’s account. Its measurement period also predates this launch, so it cannot be presented as results already accumulated by the new Sol and Luna models. Usage example

What interests me more is the effect on work that undergoes repeated revision. OpenAI allows reasoning effort to change through a specified mechanism while preserving the cache. Feature explanation A product could first organize a document with less reasoning, then devote more effort to uncertain passages. That creates an opportunity to spend on unresolved questions, provided the background information remains valid.

Suppose a team uses an agent to revise a proposal. Stable requirements and references can be reused while new feedback is added. Rearranging the entire input after every small edit could lose an otherwise usable cache. I would preserve document versions and change records in the product to reduce unnecessary retransmission. This is a design choice derived from the caching mechanism, not a measured customer benefit.

Information updates must not be delayed to protect a cache hit. If a quotation has changed, retaining the old price can make the proposal less useful even while reducing inference charges. The content should then be updated, accepting the necessary processing cost. Caching can preserve background that remains valid; it cannot decide whether information has become obsolete.

I would therefore not translate cheaper input directly into more automatic runs. A more useful comparison is spending per accepted deliverable under unchanged requirements, counting cache writes, output and failed retries together. If the bill falls but outdated information is used more often, the content-update process needs attention first. Savings must remain conditional on the result being usable.

The cover reuses the site’s GPT-6 brand illustration, not a Sol or Luna product screenshot. The new image download could not be completed because name resolution failed.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.