When a team has to process tens of thousands of tickets per day with AI, the most expensive ones are not necessarily the most difficult; rather, they are the numerous requests that seem simple but occur repeatedly. OpenAI recently expanded its GPT-6 series to include Sol and Luna precisely for such tasks. The flagship Astra continues to undertake high-difficulty tasks, Sol focuses on complex coding and proxy work, while Luna leans towards high-frequency tasks that are cost-sensitive. OpenAI announced that compared to the previous promotional prices of the GPT-5.6 series, the unit price of the two new models API has been reduced by about half. This is a change in product pricing, but it does not mean that every company's final AI budget will automatically be halved as a result.
The officially listed standard input/output prices are as follows: for every million token, the cost of GPT-6 and Sol is 2/10 US dollars, while for Luna it is 0.10/0.50 US dollars. There is a significant difference between the two prices, but we cannot simply assign all tasks to Luna based on this alone. If a task requires repeated use of tools, tracking of status across multiple files, or if the output must withstand manual review, running the lower-priced models several more times or even redoing the work could result in increased costs and time. For ordinary users, OpenAI has already provided both models as part of the corresponding payment plans for Work and Codex. Luna is also available for some free or Go users to try out on the desktop; however, the official statement also mentions that these models were not yet integrated into the regular Chat interface at that time, and that their rollout would be gradual. Confusing terms like “API available”, “Work available”, and “visible to all chat users” can easily lead to misunderstandings.
Why are there three different tiers for models of the same generation?
For some time now, discussions within the industry about models have often revolved around "who is at the top of the rankings." However, the selection of models by companies is more akin to arranging employees: some are adept at handling complex problems with unclear boundaries, some are suitable for repetitive tasks with clear rules, and others are responsible for finding a balance between speed and quality. OpenAI places Astra in the highest capability category in its official documentation, while Sol strives to make more powerful reasoning capabilities and tool usage accessible at scalable prices. Luna focuses on offering lower costs and faster daily processing. The coexistence of these three categories is due to the fact that real-world work does not come in a single level of difficulty.
OpenAI has published several tests, among which Sol achieved high scores in cross-application business process evaluations such as those conducted by AutomationBench. Luna has also made improvements over its predecessors in various coding and professional tasks. Tests are valuable in that they inform developers about which tasks are worth considering for inclusion in a candidate list. However, the prompts, tools, effort levels, and billing methods used in different tests are not consistent. The official team also notes that comparisons with some competitor models are based on public reports, and there may be differences between the research environments and actual products. Therefore, a single test result should not replace internal acceptance, nor can it be used to conclude that a new model has outperformed in all corporate scenarios.
The other side of price is caching. Long conversations and proxy tasks involve the repeated submission of the same background materials, rules, and historical steps. If the same prefixes can be cached, both the input cost and waiting time have the potential to decrease. OpenAI mentions that improvements have been made to the default cache hit rate for GPT-6, and diagnostic tools have been added for developers. There is a business issue that is easily overlooked: how much savings caching provides depends on whether the company's own workflow has reusable, stable prefixes; if each request is completely different, the savings seen in others' cases should not be directly applied. Those building systems need to consider both the cache hit rate, the total number of calls, and the quality of the final results, rather than just looking at the input unit cost.
When migrating a business, don't mistake being "smarter" for an excuse to avoid acceptance checks.
OpenAI indicates that Sol and Luna have improvements over their respective predecessors in terms of factual accuracy, coding, and computer operations. However, the internal assessment of "fewer factual errors" is based on dialogue samples where users had previously reported errors, and these dialogues do not represent the average error rate of daily requests. For sensitive tasks involving financial, medical, legal, or corporate internal data, it is still necessary to maintain source verification, access control, and manual review. While stronger models can reduce some mechanical labor, they do not automatically eliminate the need for a chain of responsibility.
Developers should be particularly vigilant about version differences that may arise during the upgrade process. The update logs of API for OpenAI also record that on September 25th, an issue with image encoding in Sol and Luna was fixed. Officials recommend that users who rely on image input to re-run evaluations and retry affected processes. Although this incident is separate from the release of the new model, it indicates that the production performance of the same model can also change after fixes are made. Teams responsible for image understanding and interface operations should not use the test scores on the release day as a permanent baseline; they should at least retain representative cases in key processes and re-test after model or platform updates.
A more prudent approach to deployment is to first stratify by task type: tasks that require in-depth judgment should be retained at the higher capability level; for tasks with clear boundaries and that can be automatically accepted, Sol should be considered; only when the quality standards are indeed met should a large number of repetitive tasks be handed over to Luna. Each layer should record the success rate, the rate of manual intervention, delays, and the total cost per task, not just the token expenditure. GPT-6, Sol, and Luna broaden the range of options available, and thus the true competition shifts from a mere performance ranking to whether a company can design its processes clearly enough.
Cover photography: Real photos of Steve Jurvetson taken by Sam Altman, Wikimedia Commons; cropped for editorial purposes; the old photos are not presented as being taken at the location of this release.












