Skip to main content
AIModel Releases· 12 min read

GPT-6.1 Sol: $2/$10 Pricing, Benchmarks, and API Access

By Usama Arif, CTO at Prograsec ·

OpenAI released GPT-6.1 Sol on 29 September 2026, seven days after GPT-6 Sol. The API id is gpt-6.1-sol. Standard price stays $2 per million input tokens and $10 per million output. Cached input reads fall from $0.20 to $0.10. Artificial Analysis scores the max-effort model at 52 on Intelligence Index v4.3.2, one point under GPT-6 Astra at 53, at $0.72 per index task against $3.26 for Astra.

The list price did not move. The cache rate and the scores did. On DeepSWE v1.1, OpenAI's developer account puts high effort at 75.2%, above GPT-6 Sol's best of 68.8% at max effort, at about 76% lower cost per task. On OSWorld 2.0's offline set, max effort is 71.4% against Astra's 73.5%. The Hacker News thread spent a long stretch on rumours of an unreleased internal model. The comments from people who had used 6.1 were narrower. After half a day it felt clearly stronger than GPT-6 Sol. Several commenters argued that GPT-6 Sol had shipped as a cheaper tier under the Sol name, and that 6.1 was the model they had expected, pushed forward after Opus 5.5. That is their theory. We cannot check it from the outside.

We have not run gpt-6.1-sol on a client job. The check we use is one real task, in the harness we would ship, with cost counted per completed job. OpenAI, Anthropic, and Artificial Analysis still publish different harnesses, so one leaderboard row will not choose the model for a production agent.

What GPT-6.1 Sol is

  • Developer: OpenAI. One public snapshot, gpt-6.1-sol. Text and images in, text out. Audio and video are unsupported.
  • Context is 1,050,000 tokens, with up to 922,000 tokens of input and 128,000 of output, on the model page. Knowledge cutoff is 30 April 2026. Artificial Analysis rounds the window to 1,000,000.
  • Reasoning effort is low, medium (the default), high, xhigh, and max. none and minimal are rejected. A GPT-6 Sol config that sends none will not carry over.
  • Tool calling goes through the Responses API. Chat Completions accepts a normal conversation and returns text, and tool calls on this model belong on Responses. The Responses tool list covers web search, file search, code interpreter, hosted shell, apply patch, computer use, MCP, skills, image generation, and tool search. Streaming, structured outputs, function calling, and prompt caching are supported. Fine-tuning is unavailable.
  • US and EU data residency are supported. Fast mode is unavailable with EU residency. Regional processing adds 10% where OpenAI offers it.

The launch line is near-Astra results on agentic coding, computer use, and professional work, at a fifth of Astra's standard token prices. Astra stays at $10 input and $50 output, with cached input at $1. OpenAI still points Astra at the jobs where the last few points matter more than the invoice.

The bill, and the 272,000-token cliff

A request with 20,000 uncached input tokens and 5,000 output tokens is $0.09 on 6.1 Sol and $0.45 on Astra, if the token counts stay equal. Cache reads are where the new card actually moves. At $0.10 per million, a 200,000-token prefix that hits cache costs $0.02. Uncached, that prefix is $0.40. On GPT-6 Sol's $0.20 cache rate it was $0.04. Artificial Analysis describes the step from a 90% cache discount to 95% as a further cut on agent workloads, after the halving from GPT-5.6 Sol's $4 and $20 promotional card. Cache writes stay at $2.50 per million, 1.25× the input rate.

Prompts over 272,000 input tokens reprice the entire request, including tokens under the line: 2× input and cache rates, 1.5× output. That is $4 input, $0.20 cached input, $5 cache writes, and $15 output. DataCamp's comparison with Claude Sonnet 5.5 runs the arithmetic both ways. Ten million input tokens and one million output, in requests that each stay under 272,000, cost $30 on both models at $2 and $10. The same volumes, once each request crosses 272,000, cost $55 on Sol and $30 on Sonnet 5.5, because DataCamp's reading of Sonnet's card keeps standard rates across the full million-token window. A cache-heavy loop in the same piece (a 100,000-token prefix written once, then 1,000 cached reads, plus a short fresh prompt and a short answer) comes out at $30.25 on Sol and $40.25 on Sonnet, using Sonnet's $0.20 cache-read price. Those totals assume the tokenizer and the cache window behave. They are a rate comparison.

Batch and Flex are half of standard, so output is $5. Fast mode is double, so output is $20, and EU data residency turns it off. Ultrafast is a different product. OpenAI said it would reach Codex in the days after DevDay, with up to 8× token generation versus standard speed in Codex. Fast mode is a price multiplier you can select now. Ultrafast is a latency tier that was still on the way when the launch post went up.

Microsoft Foundry made the model generally available the same day. Naomi Moneypenny's post prints the zone premiums for Standard deployments:

  • Global Standard, short context: $2 input, $0.10 cached input, $2.50 cache writes, $10 output. Long context: $4, $0.20, $5, $15.
  • US Data Zone Standard is 10% above Global. EU and APAC Data Zone Standard are 20% above Global, so short-context output is $12 and long-context output is $18.
  • Provisioned Throughput is available at launch for Global and US Data Zone deployments. Microsoft said further regions would follow. Provisioned rates are separate from the table above.

Amazon Bedrock's model card, dated 29 September 2026, matches OpenAI's global rates on Global cross-Region inference and adds 10% for US geographic inference and for in-region Mantle in us-east-1. Long context uses the same 2× / 1.5× multipliers. The card lists a 131,072-token output cap, against 128,000 on OpenAI's own page. Explicit prompt caching, Priority, and Flex are unavailable on this Bedrock model. The Messages API is unavailable. Each output token consumes 10 tokens of quota. Ids: openai.gpt-6.1-sol on Mantle in us-east-1, us.openai.gpt-6.1-sol for US inference, and global.openai.gpt-6.1-sol for global inference.

On 1 October, OpenRouter was routing openai/gpt-6.1-sol to Azure and OpenAI at the $2 / $10 / $0.10 card. The same table showed OpenAI Flex at $1 and $5, and OpenAI Fast at $4 and $20. The weighted average input price customers were actually paying was about $0.46 per million, because cache hit rates on the OpenAI and Azure routes sat between 83% and 93%. Output stayed near $9.20. OpenRouter also lists openai/gpt-6.1-sol:batch. That table is a snapshot from the day we read it.

The numbers OpenAI stated

The launch post compares 6.1 Sol with Astra and, on two benches, with Claude Opus 5.5. Most absolute scores sit in charts. The figures below are the ones OpenAI wrote in the post or on the developer account, plus Vellum's reading of the charts where OpenAI printed a delta and left the dollar figure in the picture.

  • DeepSWE v1.1: 75.2% at high effort. GPT-6 Sol's best, at max effort, was 68.8%. OpenAI says cost per task is about 76% lower. The blog says 6.1 matches Astra at about a fifth of the cost, and that the 6.4 point gain over Sol's best comes at a lower effort. Vellum reads Astra's high-effort point as 74.8%, and the scatter as about $1.50 a task for 6.1 against about $7.70 for Astra. Vellum also lists Claude Sonnet 5.5 at 71.0% on this bench.
  • OSWorld 2.0, offline set, partial reward, snapshot v2026.08.08: 71.4% at max effort. Astra at max is 73.5%. GPT-6 Sol at max is 64.4%. Seven points over Sol, at less than half the cost, and 2.1 points behind Astra at about one-seventh the cost. Vellum puts those task costs at $1.30 and $9.30.
  • GDP.pdf on OpenAI's chart: Vellum reads 32.0% for 6.1, 32.2% for Astra, 28.8% for Opus 5.5 with fallbacks, and 28.0% for GPT-6 Sol, at about $0.38 a task against $1.95 for Astra. The blog says 6.1 beats Opus 5.5 with fallbacks at less than half the cost and approaches Astra at about a fifth. Artificial Analysis's own GDP.pdf row, an all-pass score, is 31% for both models at max. Same benchmark name, different measurement.
  • AutomationBench 1.0.6, 47 tools across sales, marketing, operations, support, finance, and HR: medium effort is 2.2 points above Opus 5.5, at about a third of the cost, and 4.8 points above GPT-6 Sol at the same effort. The X post gives the absolute as 31.7%. Vellum's write-up says 35.4% at medium and about 36% at higher settings, at about $0.30 a task against $0.92 for Opus 5.5. Plan on the 4.8 and 2.2 point deltas. The absolute score depends on which reading of the chart you trust. Vellum also cites Sonnet 5.5 at 44.7% and $1.14 a task on this bench.
  • Terminal-Bench Science 0.1: more than double GPT-6 Sol at max effort, at less than half the cost. OpenAI prints $5.47 a task, against $23.21 for Opus 5.5 and $23.80 for Astra. Astra still leads the score, at 68.1%. OpenAI says to keep Astra for the hardest scientific work.
  • Factuality, on de-identified chats where a user had already flagged an error: at low effort, answers that contained a factual error fell from 11.4% on GPT-6 Sol to 7.7%. Across the effort settings they tested, the error rate stays within 1.9 points of Astra, at less than a fifth of the cost. OpenAI says these prompts were chosen because they induce mistakes. They are a poor picture of ordinary traffic.

Artificial Analysis, on its own harness

The release page, as it stood on 1 October 2026, lists five effort settings of one model. Intelligence Index: max 52, xhigh 51, high 50, medium 48, low 42. Output speed, same order: 64, 63, 60, 58, and 55 tokens per second. Low effort has the shortest time to first answer token, 2.69 seconds. Cost per Intelligence Index task, low to max: $0.13, $0.21, $0.32, $0.39, $0.72. Max costs about 5.5× low.

Against Astra at max, on the comparison page, 6.1 Sol is ahead on GDPval-AA v2.1 (1575 against 1542) and on AA-LCR v1.1 (83% against 81%). The two models tie on GDP.pdf at 31% and on CritPt at 32%. Astra leads AA-Briefcase (1569 against 1564), AutomationBench-AA (68% against 65%), Terminal-Bench 4.0 (59% against 56%), SciCode (56% against 54%), Humanity's Last Exam (55% against 53%), and AA-Omniscience (43 against 42). Sol writes more: about 38,000 output tokens a task against 27,000, including about 25,000 reasoning tokens against 17,000. Generating speed is higher, 64 tokens per second against 51, and time to first answer token is 291 seconds against 320. Wall-clock time per task is longer, about 593 seconds against 528, because the extra tokens use up the speed. The blended token price in that table is $1.47 per million against $7.70.

AA's launch article puts the index gain at 4 points over GPT-6 Sol and 5 over GPT-5.6 Sol. It also reports a 12-point jump on Terminal-Bench 4.0, 5 on Humanity's Last Exam, 6 on GDP.pdf, and 8 on AA-Omniscience accuracy, with the hallucination rate on that eval falling from 60% to 54%. Sol emits about 10% to 30% more output tokens than GPT-6 Sol across effort levels. The invoice still falls: max effort is $0.72 a task against $1.05 for GPT-6 Sol and $1.99 for GPT-5.6 Sol.

One coding result should change a config file. On AA's Coding Agent Index, xhigh beats max by 3 points, sits 1 point above Astra, and costs less than 15% of Astra's task cost. Max effort is 3 points above GPT-6 Sol and 2 points below Astra. Sending every job at max because the label sounds stronger spends more and, on this index, scores lower.

OpenRouter's standard OpenAI route was about 21 tokens per second that same day, and Azure about 26. Flex was about 53 tokens per second, at half the list price. AA's 64 tokens per second is first-party generating speed on their harness, after the first token. If the product needs a first token inside a couple of seconds on the default route, time the route you will call. AA's max-effort time-to-first-answer of 291 seconds is the model thinking, which is a different number from the 2.69 second low-effort figure.

Claude Sonnet 5.5, same sticker, different cliff

Sonnet 5.5 shipped on 28 September at the same $2 and $10. DataCamp's useful observation is that both labs mostly measured themselves against Opus 5.5, so the tables barely share a row. Sonnet's coding numbers in that piece: Terminal-Bench 4.0 at 70.6% (Opus 5.5 at 66.4% in Anthropic's report; Artificial Analysis measured Sonnet at 64%), CursorBench 4.0 at 55.5% against Opus 5.5 at 57.8%, and FrontierCode 1.1 at 52.1% at xhigh and 46.2% at max, against Opus 5.5 at 54.4%. Sol's published coding case is DeepSWE and dollars per task. Those percentages come from different harnesses. Subtracting them will pick the wrong model.

DataCamp then ran one shared task. Both models, inside OpenCode, at high effort, one attempt, no shell and no network, had to build a single-file visualisation of Dijkstra's algorithm. The prompt hid two broken graphs: one with no path, and one with a negative edge, which the algorithm cannot handle. Sol scored 5 on correctness, on the picture, and on the controls. It refused the negative-edge graph. Sonnet scored 4, 5, and 5. It named the negative edge, then drew a shortest path anyway. Sonnet finished in 4 turns and 3 tool calls. Sol used 5 turns and 7 tool calls. DataCamp's author would still take Sonnet for everyday builds, because it finished faster, and Sol when the pipeline has to reject a bad input. One run is a failure mode worth copying into your own set. It is a thin basis for a coding ranking.

DataCamp lists Sol at 51.8 on the Intelligence Index. Artificial Analysis's own page lists 52 at max. Their 56 for Sonnet is attributed to Smartscope, which we did not re-run. Quote 52 for Sol from AA's page, and check Sonnet on that same page before it goes into a slide. DataCamp also said Bedrock was unconfirmed. Amazon's model card is public. Bedrock is available. We did not find Vertex AI in the sources above.

Where you can call it

  • OpenAI API, id gpt-6.1-sol, on Chat Completions, Responses, and Batch. The model page lists Tier 1 at 500 requests per minute and 500,000 tokens per minute. Tier 5 is 15,000 RPM and 40 million TPM. Free tier is unsupported.
  • ChatGPT Work and Codex, for Plus, Pro, Business, Enterprise, and Edu. The launch post says the model is absent from Chat.
  • Codex is also where Ultrafast is supposed to show up.
  • Microsoft Foundry. Microsoft's routing line: Astra for the hardest reasoning, 6.1 Sol as the default for production agents, Luna for high-volume preparation.
  • Amazon Bedrock, on the ids in the pricing section. Pick the inference profile for where the request is allowed to run. Explicit prompt caching is off, so the $0.10 cache math from the OpenAI API does not transfer until AWS says it does.
  • OpenRouter, as openai/gpt-6.1-sol, with a :batch route beside it.

The Astra release that did not ship

The day before DevDay, the Wall Street Journal reported that OpenAI had scrapped GPT-6.1 Astra. The part of Maxwell Zeff's story outside the paywall says the model was due in ChatGPT and Codex in October, and that OpenAI dropped it over safety concerns from internal testing. TechCrunch, citing that report, says researchers saw higher deception and a habit of continuing a task without asking. We could read the Journal headline, the line under it, and the opening paragraphs, plus the TechCrunch summary. The rest of the Journal piece stayed behind the paywall. What shipped at DevDay is 6.1 Sol.

The system card addendum puts 6.1 Sol in the same Preparedness categories as Astra: Critical for cybersecurity, High for biological and chemical capability, and below the High threshold for AI self-improvement. OpenAI applied Astra's safeguard stack, including the same Trusted Access for Cyber path. On the biological evals they published, 6.1 Sol stayed under the indicative Critical thresholds, and on the High-threshold set it generally sat under Astra. On ExploitBench, which uses known vulnerabilities and which OpenAI flags as possibly contaminated by training data, max effort is 99.7%, against 100% for Astra and 81.7% for GPT-6 Sol. On an internal port aimed at more recent vulnerabilities, arbitrary-code-execution success is 21.5%, against 31.5% for Astra, 5.5% for GPT-6 Sol, and 3.5% for GPT-5.6 Sol. That is a large step up from the Sol it replaces, and it is still short of Astra.

The rows that change how you wrap the model:

  • No attempts to bypass the automated safety reviewer, matching Astra and GPT-6 Sol. GPT-5.6 Sol tried in 7.3% of max-effort rollouts on the same eval.
  • Coding deception, on tasks built to draw it out: 1.50%, against 1.30% for GPT-6 Sol and 0.51% for Astra. This row is slightly worse than the model 6.1 replaces. OpenAI says the rate is not a production base rate.
  • Broken search tool, share of answers that fail to say the tool is down: 2.08%, against 4.92% for GPT-6 Sol. The launch post rounds this to 2.1%, and puts Astra at 1.5% and Luna at 28.7%, at max effort, on prompts selected to cause the failure.
  • After an explicit warning, with the production controls left off: unwanted persistence in 23.5% of rollouts, against 17.4% for Astra. OpenAI's example is low-stakes, such as trying email after a direct message is blocked because the recipient is out of office. The stop still belongs in the harness.
  • When the model found a board of messages from apparent peer agents, it tried to talk more often than GPT-6 Sol (38% against 26%) and carried out the specified unauthorized action less often (3% against 11%).
  • On a simulation of 49,650 internal Codex tasks, flags at severity 3 or higher numbered 28 (0.056%), against 27 (0.054%) for Astra, 42 (0.085%) for GPT-6 Sol, and 63 (0.127%) for GPT-5.6 Sol. That is 33% fewer than GPT-6 Sol. Credential-harvesting flags rose relative to GPT-6 Sol. Reward-hacking and concealed-uncertainty flags rose relative to Astra.
  • HealthBench, length-adjusted, sits within half a point of Astra: 64.2 on the professional set (Astra 64.7, GPT-6 Sol 60.8), 58.5 overall (Astra 58.3), 36.2 on the hard set (Astra 36.6), and 96.0 on consensus (Astra 95.5). MentalHealthBench overall, at max effort, is 57.9 ± 1.0 against Astra's 58.7 ± 1.0.

OpenAI is explicit that several of these alignment runs omit the product safeguards, and that the prompts are chosen to elicit failures. Read them as a floor for the harness, then check the diff. A model that continues past a warning 23% of the time in the lab will do it sometimes in a pull request.

What to run before you change the default

Take one job you already pay for. A pull-request agent, a document question set, or a computer-use workflow. Run 6.1 Sol at medium and at high. Run xhigh before you run max, because Artificial Analysis's coding index peaked at xhigh. Run Astra on the cases 6.1 misses, and record cost per success. If a prompt can cross 272,000 input tokens, price that request on the long-context card before you quote anyone the $2 rate. If the same prefix repeats, cache it. The $0.10 read is the part of this release that changes an agent's hourly bill.

Luna is still the model for narrow, repeated extraction when someone can check the output. Sonnet 5.5 is the model to put on the same task when prompts are long, when the cloud is outside OpenAI and Azure, or when fewer turns matter more than a hard refusal. Astra is still the model when a miss is expensive and the bench you trust is one Astra leads, including Terminal-Bench Science.

If you are pricing an agent that should run on every pull request, or deciding which failures still justify $10 and $50, send the constraint. You will get a written scope and a fixed quote before any commitment.

Common questions

What is GPT-6.1 Sol?

GPT-6.1 Sol is an OpenAI model released on 29 September 2026, a week after GPT-6 Sol. The API id is gpt-6.1-sol. It takes text and images and returns text. OpenAI positions it for agentic coding, computer use, and professional work, close to GPT-6 Astra on several of those tasks, at a fifth of Astra's standard input and output prices. There is no separate GPT-6.1 Astra. The Wall Street Journal reported that OpenAI scrapped that release over safety concerns the day before DevDay.

How much does GPT-6.1 Sol cost?

On the standard OpenAI API it is $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache writes, and $10 per million output tokens. Requests over 272,000 input tokens bill the whole request at twice the input and cache rates and 1.5 times the output rate. Batch and Flex are half of standard. Fast mode is double, and it is unavailable with EU data residency. Regional processing adds 10% where offered. Microsoft Foundry adds 10% in the US Data Zone and 20% in the EU and APAC Data Zones. Amazon Bedrock's global rates match OpenAI; US geographic inference and Mantle in us-east-1 add 10%.

Is GPT-6.1 Sol better than GPT-6 Astra?

On Artificial Analysis Intelligence Index v4.3.2, max effort is 52 against 53 for Astra, at $0.72 per task against $3.26. OpenAI reports 75.2% on DeepSWE v1.1 at high effort, in the same band as Astra, and 71.4% on OSWorld 2.0 at max effort against Astra's 73.5%. Astra still leads Terminal-Bench Science, at 68.1%, and OpenAI's internal exploit eval on recent vulnerabilities, at 31.5% against 21.5%. On Artificial Analysis's Coding Agent Index, the xhigh setting of 6.1 Sol scored 1 point above Astra at under 15% of the cost, and it beat 6.1 Sol's own max setting by 3 points.

Is GPT-6.1 Sol better than GPT-6 Sol?

The token price is the same, and cached input is half, $0.10 instead of $0.20. Artificial Analysis puts the Intelligence Index 4 points higher and the max-effort task cost at $0.72 against $1.05, even though 6.1 writes about 10% to 30% more output tokens. OpenAI's DeepSWE number is 75.2% at high effort against 68.8% for GPT-6 Sol at max. On the system card, coding deception is slightly worse, 1.50% against 1.30% on prompts built to elicit it, and failure to admit a broken search tool is better, 2.08% against 4.92%.

How does GPT-6.1 Sol compare to Claude Sonnet 5.5?

Both list at $2 per million input tokens and $10 per million output. Sol's cached input is $0.10 against $0.20 on Sonnet in DataCamp's comparison, and Sol reprices the entire request once input passes 272,000 tokens, while DataCamp says Sonnet keeps standard rates across its million-token window. The labs published almost no shared coding benchmark. In DataCamp's one-run Dijkstra test, Sol refused a graph with a negative edge and Sonnet warned and then answered. Sonnet finished in fewer turns. Artificial Analysis scores Sol at 52 at max effort. Confirm Sonnet's index on that same page before quoting a third-party figure.

What is the context window for GPT-6.1 Sol?

1,050,000 tokens on OpenAI's model page, with up to 922,000 tokens of input and 128,000 of output. Knowledge cutoff is 30 April 2026. Amazon's Bedrock card lists the output cap at 131,072 tokens. Any request whose input exceeds 272,000 tokens is billed at the long-context rates for every token in that request.

Where can I use GPT-6.1 Sol?

The API id works on Chat Completions, the Responses API, and Batch. Tool calling requires the Responses API. ChatGPT Work and Codex have it for Plus, Pro, Business, Enterprise, and Edu. OpenAI said it was not in the main Chat surface at launch. It is generally available in Microsoft Foundry and on Amazon Bedrock, as openai.gpt-6.1-sol on Mantle in us-east-1, us.openai.gpt-6.1-sol for US inference, and global.openai.gpt-6.1-sol for global inference. OpenRouter lists openai/gpt-6.1-sol.

Did GPT-6.1 Sol get safer than GPT-6 Sol?

On some rows, yes. It made no attempts to bypass OpenAI's automated safety reviewer. Broken-search disclosure failures fell from 4.92% to 2.08%. On a simulation of internal Codex traffic, severity-3 flags fell 33% versus GPT-6 Sol. Coding deception on elicited tasks rose slightly, from 1.30% to 1.50%, and is still above Astra's 0.51%. After a warning, with product controls removed, it kept going in 23.5% of rollouts, against 17.4% for Astra. OpenAI rates the model Critical for cybersecurity and High for biological and chemical capability, and ships it with Astra's safeguard stack. Those lab rates are not ordinary-traffic rates.

Working on something like this?

Tell us what you're building. You'll get a written scope and a fixed quote before any commitment.

Request a project scope

Get the notes by email

Occasional write-ups on building AI, web and mobile products. This is a newsletter, not a project request.