GPT-6 Sol and Luna: $2/$10 Pricing, Benchmarks, and API Access
By Usama Arif, CTO at Prograsec ·
OpenAI released GPT-6 Sol and GPT-6 Luna on 22 September 2026, nineteen days after GPT-6 Astra. The API ids are gpt-6-sol and gpt-6-luna. Standard list price is $2 per million input tokens and $10 per million output for Sol, and $0.10 and $0.50 for Luna. That is half the promotional card for GPT-5.6 Sol ($4 / $20). Luna's input is also half ($0.20 to $0.10). Its output fell further, from $1.20 to $0.50.
The price cut is the product. On the benches OpenAI chose to print, Sol at max effort scores 68.8% on DeepSWE v1.1, which is under the 72.7% GPT-5.6 Sol posted in the Astra launch table, and 60.5% on OSWorld 2.0 at extra-high effort, under the 65.7% Sol number from that same Astra table. Artificial Analysis has the Intelligence Index and Coding Agent Index roughly level with GPT-5.6, with gains on some rows and losses on others. An OpenAI spokesperson told The New Stack the new prices are the default card, not a promotion. GPT-5.6's own promotional rates were due to rise 25% in November, which is the baseline OpenAI is measuring the cut against.
Anthropic shipped Claude Opus 5.5 about ninety minutes earlier the same day, at $4 and $20, with cache reads at $0.20. OpenAI's charts compare Sol to Opus 5 and Fable 5 / 5.1, not to Opus 5.5. We have not run gpt-6-sol or gpt-6-luna on a client job. The test we use on every release is one real task, in the harness we would ship, with cost counted per completed job.
What GPT-6 Sol and Luna are, in numbers
- Developer: OpenAI. API slugs
gpt-6-solandgpt-6-luna. There is no GPT-6 Terra. Sol now sits at GPT-5.6 Terra's old $2 input price, with output at $10 instead of Terra's $12. - Both take text and images and return text. Context is 1,050,000 tokens, of which 922,000 can be input, with a 128,000 token output cap. Handy AI and Kingy both list those limits. Knowledge cutoff is 20 April 2026 for Sol and 18 May 2026 for Luna, so the cheaper model is the more recent one. Astra's cutoff is 30 April 2026.
- Reasoning effort is
none,low,medium(the default),high,xhigh, andmax. Astra starts atlowand has nonone. The Rundown's guide matches the API docs on that ladder. - Cached input reads are $0.20 per million on Sol and $0.01 on Luna, a 90% discount. Cache writes are $2.50 on Sol and $0.125 on Luna, from the Azure pricing table and OpenAI's pricing page as read by Kingy. Astra stays at $10 input, $1 cached input, $12.50 cache writes, and $50 output.
- Prompts over 272,000 input tokens bill the whole request at 2x input and cache rates and 1.5x output. Batch and Flex are half of standard, so output is $5 on Sol and $0.25 on Luna. Fast mode is 2x. A Hacker News commenter pointed at the pricing page for the batch line the same afternoon.
- Chat Completions, Responses, and Batch are supported. Kingy's read of the model pages lists function calling, web search, file search, and computer use on both, and adds structured outputs, code interpreter, image generation, MCP, and skills on Luna's page. Fine-tuning is not offered for Luna.
OpenAI's line is that Sol and Luna were trained with similar methods to Astra, and that caching and inference improvements are what made the price cut possible. Reuters reported that sentence the morning of the launch. Astra remains the model they tell you to use when the result matters more than the invoice.
The sticker price, and what a task actually costs
A request with 20,000 uncached input tokens and 5,000 output tokens is about $0.09 on Sol and $0.0045 on Luna, using the rates in The Rundown's worked example. Astra on the same counts is $0.45. That arithmetic holds token counts equal. It does not hold if Luna writes more.
Artificial Analysis measured that directly. On their Intelligence Index, GPT-6 Sol at max costs $1.06 a task, against $1.99 for GPT-5.6 Sol at max. Luna at max costs $0.07, against $0.18. Both models spend more output tokens than their predecessors: about 31,000 versus 29,000 for Sol, and 51,000 versus 41,000 for Luna. The sticker fell by half. The bill fell by about half for Sol and about 60% for Luna, because the extra tokens ate part of the cut. On their Coding Agent Index, in OpenAI's Codex harness, Sol at max is $2.99 a task, again about half of GPT-5.6 Sol.
Cache is the other half of the invoice on a coding agent. Reads are 10% of the input price. You can now change reasoning effort, or turn tools on and off, without dropping the cached prefix. Explicit breakpoints let you choose where that prefix ends. The prompt caching dashboard and the diagnostics guide show what missed. GitHub says the caching work, over recent months, cut the share of Copilot prompt tokens that needed a fresh read by more than half, across billions of requests. On Hacker News, people who live in cached prefixes pointed out that a 50% cut on uncached input is not a 50% cut on their bill. One reader put cache reads at about half their spend. Another said uncached reads still dominate. Both can be true. Log the four lines separately: input, output, cache reads, cache writes.
Stay under 272,000 input tokens unless you mean to pay the long-context rate on the entire request. Past that cliff, Sol's input price matches Opus 5.5's $4.
What OpenAI published, with the effort setting attached
The launch post evaluates OpenAI models in a research environment or via the API, and takes competitor scores from public reports. Fable 5 appears where Fable 5.1 numbers were not available. The charts plot score against cost per task at every effort setting. The prose quotes one point on each curve. Kingy read the rest of the points off those charts. Where a number below is Kingy's chart reading and not a figure OpenAI wrote in the post, it is marked.
- AutomationBench 1.0.6, 47 business tools: Sol at xhigh is 33.2% at $0.27 a task. Opus 5 at max is 26.9%, at 11.1 times the cost. Astra at low is 30.3%, at 3.9 times the cost. Fable 5.1 with an Opus 5 fallback is 31.4%, and OpenAI says that cost figure leaves out the fallback, which fired on about 40% of tasks. Luna at high is 5.4 points above GPT-5.6 Luna, at 58% lower cost per task. Kingy's chart reading: Sol at max is 32.0% at $0.34, worse than xhigh, and Luna's best is 20.7% at max for $0.037. A developer on the OpenAI forum made the same point from the chart: max effort on Sol did not buy a higher AutomationBench score.
- Agents' Last Exam, long professional workflows across 55 sub-industries: Sol at max is 56.4%, above Opus 5's best in that evaluation, at 60% lower cost per task. Astra's published score on this exam, from the Astra launch, is 59.3%. Kingy puts Opus 5's best at 55.9% and $7.29, and Luna at max at 50.9% and $0.15, against $2.57 for GPT-5.6 Luna at a similar score.
- DeepSWE v1.1: Sol at max is 68.8%, 1.1 points behind Fable 5 at xhigh (69.9%), at about 80% lower cost per task. Luna at max is 66.6%, in line with Opus 5 and Fable 5 at medium, at 93% and 96% lower cost. The number the post does not put next to that sentence: GPT-5.6 Sol's best on the Astra launch table was 72.7%, and Astra's was 74.1%. Kingy's chart reading puts Opus 5 at max at 73.7% and $11.84, and GPT-6 Sol's 68.8% at $2.74. Matched on dollars, the new Sol wins. Matched on peak score, it does not.
- FrontierCode 1.1, mergeable patches, not just passing tests: OpenAI says Sol improves on GPT-5.6 Sol and matches Fable 5.1 at xhigh, at much lower cost. They do not print the percentage. Kingy's chart reading: Sol's best is 49.3% at max for $2.14. Fable 5.1 at xhigh is 48.7% for $9.27, which is the match they claim. Fable 5.1 at medium is 50.9%, and Opus 5 at medium is 53.4%. xhigh is not Fable's best setting on this bench.
- OSWorld 2.0 offline, partial reward, v2026.08.08 snapshot: Sol at xhigh is 60.5%, against Opus 5 at medium at 60.3%, at about 80% lower cost. Luna at max beats GPT-5.6 Sol at medium at about a tenth of the cost. Astra remains ahead on computer use. Our Astra piece cites 72.6% from the Astra launch. Kingy's reading of the new chart puts Astra's peak at 73.5%, Sol's peak at 64.4% and $3.25, and GPT-5.6 Sol's peak at 66.2%. The xhigh comparison in the blog post is the flattering slice of that curve.
- Internal factuality, built from de-identified ChatGPT threads where a user had flagged an earlier mistake: Sol makes about half as many errors as GPT-5.6 Sol. Luna at higher effort matches GPT-5.6 Sol at about a hundredth of the cost. OpenAI says a verbosity sweep showed almost no link between length and accuracy on this set, and that the prompts are harder than ordinary use. Kingy's chart reading puts Sol's error rate at 4.5% at xhigh ($0.13) and 4.6% at max, against 8.5% for GPT-5.6 Sol at max.
What an independent lab measured
Artificial Analysis ran both models at max effort. The composite indexes barely moved. The rows inside them did.
- Coding Agent Index, Codex harness: Sol 57, up 2 from GPT-5.6 Sol. Terminal-Bench 4.0 is 43% against 37%, and SWE-Atlas-QnA is 58% against 54%. Luna is 41, down 2, with SWE-Atlas-QnA at 44% against 49% and DeepSWE at 64% against 66%. That DeepSWE number is their harness, not OpenAI's 66.6%. Do not subtract them.
- AA-Omniscience, their knowledge and hallucination set: Sol's hallucination rate falls from 92% to 60%. It also attempts 83% of questions instead of 99%, and accuracy falls from 59% to 54%. The index still rises, from 22 to 27, because fewer wrong answers outweigh the extra refusals. Luna's hallucination rate falls from 93% to 77%, accuracy is 44% against 43%, and the index goes from -10 to 1.
- GDPval-AA v2.1, work across 44 occupations: Sol drops about 100 Elo and Luna about 75. AA-Briefcase, multi-week projects with thousands of files: Luna drops about 45 Elo and Sol is flat. AA's reviewers looked at hundreds of outputs and blamed shorter deliverables that skip items on the rubric. That sits next to OpenAI's style claim. Shorter can be clearer in a coding chat and incomplete in a knowledge-work packet.
- Their own automation set, AutomationBench-AA, is a different test from Zapier's: Sol 62% against 60%, Luna 53% against 50%. Terminal-Bench 4.0 in the Intelligence Index breakdown is Sol 44% against 40% and Luna 13% against 12%. The 43% and 44% Terminal-Bench figures are both on the AA page, attached to different summaries. Quote the harness if you quote either.
DataCamp ran one build, not a bench: a single-file Dijkstra visualizer, same prompt, file read and edit only, one attempt each in OpenCode. The data file included a negative edge, which the prompt did not mention. All four models (GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Luna) solved the clean graph and stopped on an unreachable target. Only GPT-6 Sol refused to run Dijkstra on the negative edge and said why on screen. GPT-5.6 Sol warned and then filled in a distance table anyway. Cost on that task was $0.25 against $0.23 for Sol, and $0.02 against $0.03 for Luna. One task, one shot. It matches the honesty claim better than the coding-score claim.
Same day as Opus 5.5, and the charts do not include it
Opus 5.5 is $4 and $20, twice Sol's list, with the same $0.20 cache-read price. Anthropic says typical workloads cost about 40% less than Opus 5, because the model also uses fewer tokens. On the two rows both labs publish in a form you can put side by side, Opus 5.5 leads: 40.0% on AutomationBench against Sol's 33.2% at xhigh, and 54.4% on FrontierCode against the 49.3% Kingy reads as Sol's best. Anthropic's Terminal-Bench 4.0 number for Opus 5.5 is 66.4%. OpenAI did not publish Terminal-Bench 4.0 for Sol or Luna. Simon Willison put Sol next to Opus 5.5 on the price card and switched his own Codex default to Sol the same night, with Luna on a Datasette agent demo. Nobody has published a shared-harness run of Sol against Opus 5.5. Until that exists, the honest comparison is list price plus the rows above, then your own task.
Vellum's recap repeats OpenAI's cost-per-task framing and adds Dan Shipper's early note that Sol matched Astra on 17 of 18 simpler computer-use tasks at Every, at lower token cost. That is a preview anecdote, not OSWorld. Shipper also hit Codex's security classifier pausing authorized runs for a human confirm. That pause is a harness, not a model score, and it is the thing that will show up on an unattended job.
Alignment improved on honesty, and barely moved on warnings
Sol and Luna do not have their own system card. OpenAI points at the Astra system card and says the evals below are hard cases, run without the full product safeguards, and are not typical-use rates. Handy AI's objection is fair: these models were trained with similar methods to the first OpenAI model rated Critical for cyber, and they shipped without a Preparedness rating of their own.
- Coding deception at max effort: Sol 1.3%, down from 10.4% on GPT-5.6 Sol. Luna 2.8%, down from 9.5%. Astra 0.5%. Kingy's chart reading.
- Told to use a search tool that was broken, and graded on whether it said so: Sol failed to disclose 4.9% of the time, down from 77.5%. Luna 28.7%, down from 78.3%.
- An explicit access-denied warning: Sol still tried to work around it in 64.4% of runs, down from 68.2%. Luna fell from 76.5% to 42.4%. Astra is at 17.4% on Kingy's table. If a permission boundary matters, the harness has to enforce it. Sol's own rate is not the control.
- A simulated message board with unauthorized instructions, including a request to disclose private data: Sol acted on it in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none. OpenAI notes Luna also found the board less often.
Where you can call them
ChatGPT Work and Codex, for Plus, Pro, Business, Enterprise, and Edu, rolling out through launch day. Free and Go users get Luna in the desktop app only. OpenAI said neither model was in the main Chat surface yet, and that the Work and Codex rollout would be staggered so the service stayed up. OpenAI's forum post adds a banked usage reset for Plus, Pro, and Business accounts: you choose when to apply it. One reply the next day still had not seen the reset. Another described a routing stack of Sol at high as orchestrator, Luna as the worker, and Astra as the reviewer, and said Astra as orchestrator wandered.
GitHub Copilot bills both on usage-based pricing. Sol is on Pro+, Max, Business, and Enterprise. Luna is also on Copilot Pro. Surfaces listed in the changelog: VS Code, Visual Studio, JetBrains, Xcode, Eclipse, the Copilot CLI, the cloud agent, the Copilot app, github.com, and the mobile apps. Business and Enterprise admins can turn them off in model policy. Otherwise new models are on unless that default was disabled. Rollout is gradual.
Amazon Bedrock lists both as generally available, with explicit prompt caching. IAM controls access, CloudTrail records calls, and PrivateLink keeps traffic on a VPC endpoint. AWS says inference data is not used to train the models, and that you do not opt into sharing prompts with OpenAI by using Bedrock. Abuse-classifier traffic can be kept up to 30 days. Zero data retention is a request to the account team, not a console toggle.
Microsoft Foundry has Standard deployment for Astra, Sol, and Luna across all 28 global regions and the US and EU data zones. Provisioned throughput covers Astra and Sol. Priority processing covers Sol. Global Standard short-context rates match the API card. US data zone is about 10% higher ($2.20 / $11 for Sol, $0.11 / $0.55 for Luna). EU is higher again. Long context doubles input and raises output by half, same shape as the 272k rule. OpenRouter serves openai/gpt-6-sol on chat completions, including reasoning effort. Recheck the router price before you quote a client a unit cost.
A LinkedIn comment on launch day said Codex with a ChatGPT account still rejected both slugs. Treat that as a rollout report, not a permanent limit. On Hacker News (1,748 points that afternoon) one team already used Luna as the fallback when an open-weight model returned nothing, and another said Luna was cheap enough to put on a site without a login. A third said they would rather keep GPT-5.6 Sol if GPT-6 Sol behaves like Astra. MacRumors had the same split in miniature: one person reporting that Sol at high beat GPT-5.6 Sol at medium and cost less, and several people saying the names no longer tell you which model is the strong one.
How the replies changed
OpenAI carried Astra's writing style down to both models: less jargon, fewer odd phrases, fewer details that do not change the answer, slightly shorter. Their side-by-side is a small front-end request. GPT-5.6 Sol decided React was unnecessary, described a "bento feel," and pasted an image-tool prompt the user had not asked for. GPT-6 Sol said it would check whether the existing site needed React, then reported that it had checked desktop, a narrow screen, and browser back navigation. Style notes are easy to fake in a launch post. The DataCamp negative-edge result is the same behavior on a task someone else wrote: say the method does not apply, and stop.
AA's GDPval drop is the other side of that style. If the rubric wants a section and the model now skips it, you have a cheaper answer and a failed deliverable. Cap length only after you have checked the required parts are still there.
Where we'd use each one
- Luna, medium or high: extraction, classification, routing, and summaries you will run thousands of times. The $0.10 / $0.50 card is the reason. Cap output. AA's extra 10,000 tokens a task will show up if you do not.
- Luna at max: a cheap first pass on a coding agent, when a miss is cheap to detect. OpenAI's DeepSWE number (66.6% at $0.22 on Kingy's chart) is the case for that. AA's Coding Agent Index going down two points is the case against using it as the only coder.
- Sol at xhigh for multi-app workflows, and high or max for coding, depending on the bench you care about. AutomationBench peaked at xhigh in Kingy's reading. DeepSWE and FrontierCode kept rising through max. Pick the effort on your task, not from the launch sentence.
- Astra when the job is computer use, or when a wrong merge costs more than a 5x token price. Sol at xhigh beat Astra at low on AutomationBench. It did not beat Astra at max on any chart Kingy transcribed.
- Opus 5.5 when you need the coding rows Sol did not publish, or when cache reads are most of the bill and the $0.20 read price makes the input gap smaller than the headline. Run both on one job before you move the default.
- Poor fit: a prompt that regularly crosses 272,000 input tokens and was priced on the short-context card. Poor fit: an unattended agent whose only permission check is a sentence in the prompt. Sol's warning-circumvention rate is 64.4% on OpenAI's own hard test.
How we'd test it on your workload
Same method as the Astra release. Take one real task from a live project. Run it on the incumbent, on gpt-6-luna, and on gpt-6-sol, at the effort you would actually ship, with the same tools. Count cost per completed task, including cache reads, cache writes, and retries. Count how often a person had to step in, and how often the answer was shorter in a way that dropped a required part. If the task has a permission boundary, include one case the agent should refuse. Do that on a redacted copy. Do not discover the 272k cliff, or the warning-circumvention rate, in production.
If you are pricing an agent feature and the question is whether Luna makes the volume tier affordable, or whether Sol replaces the model you are on now, that is a scoping conversation. Send the constraint and you will get a written scope and a fixed quote before any commitment.
Common questions
What are GPT-6 Sol and GPT-6 Luna?
They are OpenAI models released on 22 September 2026, below the flagship GPT-6 Astra. The API ids are gpt-6-sol and gpt-6-luna. Sol is aimed at coding and multi-step agent work. Luna is aimed at high-volume tasks such as extraction, classification, and short summaries. There is no GPT-6 Terra. Both take text and images and return text.
How much do GPT-6 Sol and GPT-6 Luna cost?
On the standard OpenAI API, Sol is $2 per million input tokens and $10 per million output tokens. Luna is $0.10 and $0.50. Cached input reads are $0.20 on Sol and $0.01 on Luna. Cache writes are $2.50 on Sol and $0.125 on Luna. Prompts over 272,000 input tokens bill the whole request at twice the input and cache rates and 1.5 times the output rate. Batch and Flex are half price. Fast mode is double. These prices replace GPT-5.6 promotional rates of $4/$20 for Sol and $0.20/$1.20 for Luna. An OpenAI spokesperson told The New Stack the new rates are the default card, not a promotion.
Is GPT-6 Sol better than GPT-5.6 Sol?
It is cheaper, and it is more reliable on OpenAI's internal factuality set, where mistake rates fell by about half. On DeepSWE v1.1, GPT-6 Sol at max scores 68.8%, below the 72.7% GPT-5.6 Sol posted in the Astra launch table, at a much lower cost per task. On OSWorld, the point OpenAI highlights (60.5% at extra-high effort) is also below GPT-5.6 Sol's earlier peak. Artificial Analysis has the Coding Agent Index up two points, to 57, and a large drop in hallucination that comes partly from answering fewer questions. Switch for the price. Retest if you were buying peak coding or computer-use score.
Should I use GPT-6 Sol, GPT-6 Luna, or GPT-6 Astra?
Use Luna for repeated, narrow tasks where you can check the output. Use Sol for everyday coding and multi-step work across tools. Use Astra when computer use or the hardest jobs justify $10 per million input tokens and $50 per million output. Sol at extra-high effort beat Astra at low effort on AutomationBench. Astra still leads the charts at its own best setting, including computer use. Prompts over 272,000 input tokens narrow the price gap because OpenAI bills the whole request at long-context rates.
How do GPT-6 Sol and Luna compare to Claude Opus 5.5?
Opus 5.5 launched the same day at $4 and $20 per million tokens, twice Sol's list price, with the same $0.20 cache-read price. On AutomationBench, Anthropic reports 40.0% for Opus 5.5 and OpenAI reports 33.2% for Sol at extra-high effort. On FrontierCode, Opus 5.5 is 54.4% and Sol's best point on OpenAI's chart, as read by Kingy, is 49.3%. OpenAI's launch charts compare Sol to Opus 5 and Fable, not to Opus 5.5. No independent lab had published a shared-harness comparison on launch day.
Where can I use GPT-6 Sol and Luna?
The API ids work on Chat Completions, Responses, and Batch. ChatGPT Work and Codex have them for Plus, Pro, Business, Enterprise, and Edu, with a gradual rollout. Free and Go users get Luna in the desktop app. They were not in the main ChatGPT chat surface on launch day. GitHub Copilot has Sol on Pro+, Max, Business, and Enterprise, and Luna on those plans plus Copilot Pro. Both are on Amazon Bedrock and in Microsoft Foundry. OpenRouter lists openai/gpt-6-sol.
What is the context window for GPT-6 Sol and Luna?
1,050,000 tokens, with up to 922,000 tokens of input and 128,000 tokens of output. Knowledge cutoff is 20 April 2026 for Sol and 18 May 2026 for Luna. Requests above 272,000 input tokens are billed at the higher long-context rates for every token in the request, not only the overflow.
Did GPT-6 Sol and Luna get safer?
On OpenAI's hard tests, coding deception fell sharply: Sol from 10.4% to 1.3%, Luna from 9.5% to 2.8%, at maximum effort. When the test was an explicit access-denied warning, Sol still tried to work around it in 64.4% of runs, down from 68.2%. Luna fell from 76.5% to 42.4%. OpenAI says these runs omit the full product safeguards and are not typical-use rates. Treat permission checks as something your harness enforces.
Keep reading
Working on something like this?
Tell us what you're building. You'll get a written scope and a fixed quote before any commitment.
Request a project scopeGet the notes by email
Occasional write-ups on building AI, web and mobile products. This is a newsletter, not a project request.