Skip to main content
AIModel Releases· 11 min read

Grok 4.6: Tied with GPT-5.6 Sol, Same $2/$6 as 4.5

By Usama Arif, CTO at Prograsec ·

SpaceXAI released Grok 4.6 on 12 August 2026, five weeks after Grok 4.5. The model slug is grok-4.6. The headline price did not move: $2 per million input tokens and $6 per million output, the same as 4.5. The score did. It lands at 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and five points above Grok 4.5 High.

This is a day-old release, so what follows is the launch data, the first independent numbers, and the API details that will show up on an invoice. It is not our own eval on a client workload. We'll run that the same way we always do: one real task, same harness, cost per completed job.

Two facts decide whether this is an upgrade or a press release. The first is that the jump over 4.5 is real on long agent runs and uneven on coding benches, so "tied with Sol" is a composite, not a sweep. The second is a pricing change most launch posts skip: cache hits went from $0.30 to $0.50 per million, and any prompt that crosses 200,000 tokens doubles the whole request.

What Grok 4.6 is, in numbers

  • Model ID grok-4.6. Text and image in, text out. Knowledge cutoff 1 February 2026.
  • 500,000-token context window, unchanged from Grok 4.5. No published text output cap.
  • API list price below 200,000 prompt tokens: $2.00 input, $0.50 cached input, $6.00 output per million tokens.
  • At or above 200,000 prompt tokens: $4.00 / $1.00 / $12.00, billed at the higher rate for every token in that request, not only the overflow.
  • A fast / priority tier at 2x those rates. Cursor prices the Fast variant at $4 / $1 / $12.
  • Four reasoning efforts: low, medium, high (default), and xhigh.
  • Tools on the SpaceXAI API: function calling, web search, X search, and code execution. Cursor gives it the full agent tool set.
  • Available day one in Cursor, Grok Build, the SpaceXAI API, and through OpenRouter, Vercel and Cloudflare.

Cursor is treating 4.6 as a first-party model in the same usage pool as Grok 4.5 and Composer 2.5. For the first week from 12 August they are running a 50% on-demand discount and SpaceXAI is doubling included usage in Cursor and Grok Build. After that week, budget on the rate card above.

One availability note, because we flagged the opposite for 4.5. Cursor's help docs say Grok 4.6 is offered in every country where Cursor ships models, including the EU. The API still warns that access can vary by geography, so confirm the endpoint for your region before you promise a client it will be there.

The benchmarks, without the sweep

SpaceXAI's own table, with third-party scores taken as the best published figure for each rival, is the right place to start and the wrong place to stop. Best score in each row in the launch post is bolded; I'm listing the four models they compared.

  • AA Intelligence Index: Grok 4.6 at 61, tied with GPT-5.6 Sol Max, behind Claude Fable 5 Max (62) and Claude Opus 5 (63). Grok 4.5 High was 56.
  • GDPVal-AA v2: 1753 Elo, ahead of Fable 5 Max (1741) and Sol Max (1728). Grok 4.5 was 1526. Artificial Analysis says the gap to Fable sits inside overlapping confidence intervals.
  • CursorBench v3.2: 69.9%, up from 66.7%. Ahead of Sol Max (67.2%), behind Fable 5 Max (70.5%). Cursor has previously said an earlier snapshot of their codebase was in Grok's training data, which still applies to how you read this particular bench.
  • DeepSWE v1.1: 65.9%, up from 54%. Sol Max 73%, Fable 5 Max 70%. This is the row that still loses.
  • FrontierCode v1.1 Extended: 61.3%, a narrow lead over Sol Max (60.6%), behind Fable (63.6%).
  • APEX-Agents: 57.5%, up from 47.1%. Nudges past Sol Max (56.7%), behind Fable (59.2%).
  • Terminal-Bench v3.0: 26%, up from 15.7%. Sol Max 34.6%, Fable 5 Max 34.1%. Large gain, still last of the four.
  • AA-Briefcase: 1577 Elo, roughly Fable-tier (1574) and ahead of Sol Max (1502). Grok 4.5 was 1313.

Read Terminal-Bench with the version attached. Artificial Analysis also reports 88.4% on Terminal-Bench v2.1, which is a different test from the v3.0 score of 26% in the launch table. Mixing those two numbers is how a model looks both elite and mid-pack in the same paragraph.

The honest summary: 4.6 is clearly better than 4.5. Against Sol and Fable it wins some agent and knowledge-work rows, trails on DeepSWE and Terminal-Bench v3.0, and ties Sol on the composite index. Cognition put it in Devin the same day and described the same shape: past Sol, behind Opus 5 and Fable 5.

When we compared Opus 5, Grok 4.5 and GPT-5.6 Sol in July, Grok's case was price and speed on non-visual work. That case is stronger now on capability. It is not a reason to delete the other two from a routing table.

Token price is not task cost

The $2/$6 card is more than 60% below Opus 5 ($5/$25) and Sol ($5/$30) on output, which is where reasoning-heavy work spends. That comparison is real and it is incomplete.

Artificial Analysis measured $0.84 per Intelligence Index task and put 4.6 on their intelligence-versus-cost Pareto frontier. On AA-Briefcase, the more useful figure is turns: about 53 turns and 0.5 billion input tokens on average, against about 103 turns and 2.0 billion input tokens for Claude Opus 5 Max. Half the turns and a quarter of the input is a bigger saving than the rate card, if it survives contact with your harness.

The catch is cache. Grok 4.5 billed cache hits at $0.30 per million. Grok 4.6 bills them at $0.50. On a long coding agent, cache reads are often most of the bill. Developers on Hacker News noticed this within hours, and they are right to. A model that finishes in fewer turns can still cost more per session if every reread of the prompt got 67% more expensive.

The 200,000-token cliff is the other one. Cross it and the entire request, not the tail, moves to $4/$1/$12. A 199k prompt and a 201k prompt are different products. If you stuff a repo into context "because the window is 500k," you can double the job without noticing.

The rate card is $2/$6. The invoice on a long agent loop is cache reads, retries, and whether the model finishes in 50 turns or 100.

Same 1.5T model, longer post-training

Grok 4.6 is not a bigger brain. Musk described it before launch as the 1.5-trillion-parameter model with better supervised fine-tuning and reinforcement learning, the same foundation as 4.5. The 2.1T step is labelled Grok 4.7 and is still ahead. Treat that as an executive's tweet until a model card says otherwise, but the launch write-up matches the shape: a longer supplemental run, an improved optimiser, and SFT trajectories regenerated by Grok 4.5 across STEM, software engineering and knowledge work, with bad traces filtered by model checks.

The RL environments are the tell. SpaceXAI trained on general coding and knowledge work, then on narrower setups for kernel optimisation, web development and computer-aided design. The New Stack's read is that the model was rewarded for finishing the larger task rather than emitting a plausible block of code. That is the behaviour they claim showed up later: more self-checks on long trajectories, and a stronger first pass on visual and interactive apps.

Those are company observations. They line up with what Cursor's Eric Zakariasson reported after using a pre-release build as a daily driver, which is the closest thing to a hands-on note we have on day one.

How to prompt it, from people who already have

Zakariasson's field notes are more useful than the benchmark slide. He found that "work very hard" and similar pep-talk phrasing did almost nothing. Prompt length did, and not in the direction most teams assume. A long spec buys specificity when you already know the product. A short prompt hands taste to the model, and on 4.6 that taste was good enough that a three-sentence brief and a two-page spec produced nearly the same spreadsheet app.

The line that actually changed the result was an instruction to open the running app, click through real user paths, check that nested formulas evaluated, and fix what it found. Same idea on a 3D scene: "improve the textures" went nowhere; "capture the current frame, list what's wrong, then fix only those things" worked. The model will keep going without being told to. What it will not invent for you is the definition of done.

  • Write what done looks like, including how to verify it. The verification loop is the highest-leverage sentence in the prompt.
  • Keep the first prompt short unless you already have a spec. Add constraints when the first pass is wrong, rather than front-loading every preference.
  • Do not add "keep going until it's finished." SpaceXAI trained this one to stay on long tasks; extra urging mostly burns tokens.
  • Give it a way to look. A DOM and a screenshot are easy. Video, 3D and physics are not, and you become the checker unless you capture frames or logs.
  • Set prompt_cache_key on the Responses API (or the x-grok-conv-id header on Chat Completions). SpaceXAI is blunt about this: without it, multi-turn agent loops often miss the cache and you pay full input price.

One API gotcha from the Hacker News thread on launch day: some callers saw a default Grok system prompt prepended to requests, including a line about not discussing those guidelines, which then fought their own system prompt. If the model starts refusing to talk about instructions, that is the first thing to check, not your wording.

Where we'd use it, and where we wouldn't

For high-volume agentic work where cost per completed task decides whether the feature ships, 4.6 is now the Grok you actually try, not 4.5. The AA turn counts and the $2/$6 card are the reason, provided you stay under the 200k cliff and you pin a cache key. That is the slot we previously gave 4.5 in the July comparison.

For a first pass on an interactive or visual prototype, the launch claim is that 4.6 is the first Grok worth starting with rather than discarding. Zakariasson's Age of Empires recreation is an anecdote, not a bench, but 4.5 came back with a flat prototype and 4.6 came back with an isometric world and a HUD. If your product is a UI, that first-pass gap is the whole point of trying it.

We would not make it the only model in a stack that has to win DeepSWE-shaped repo work, and we would not pick it for Terminal-Bench v3.0-shaped jobs while Sol and Fable still lead that row by eight points. We also would not put it on a regulated, customer-facing surface just because the coding scores moved. Grok's consumer brand has a public safety history that procurement teams will treat as a vendor-risk question, separate from grok-4.6's evals. That is a client conversation, in writing, before anyone argues about Elo.

  • Good fit: long agent loops in Cursor or Grok Build, high-volume knowledge work, and product-idea-to-working-first-version sessions where you will iterate in the loop.
  • Good fit: anything that was already on Grok 4.5 for price, now that the capability gap to Sol has closed on the composite index.
  • Poor fit: a single-model stack with no escalation, terminal-heavy SWE where Sol still leads, and any workflow that regularly crosses 200k prompt tokens without a compaction step.
  • Poor fit: buyers who cannot accept the Grok brand on a customer-facing or compliance-sensitive path, regardless of the rate card.

Kimi K3 is the other comparison people will search. Moonshot's open-weight model was sitting near 60 on the same index at $3/$15. Grok 4.6 is now ahead on that score at a lower output price, with a 500k window against K3's 1M. If you need open weights or a million-token context, K3 still has a job. If you need a hosted coding agent at $6 output, 4.6 is the one that moved.

How we'd test it on your workload

Same method as every release. Take one real task from a live project, run it against your incumbent and against grok-4.6 with the same harness at high effort, and count three things: cost per completed task (including cache), how often a person had to step in, and how many failures were recoverable without a human. If the task is visual, add a fourth: whether the first pass is something you would show a client or something you would throw away.

If you're weighing an AI feature and want to know whether this pricing changes what's affordable, or whether this vendor should touch your data at all, that's a scoping conversation we're glad to have. Send us the constraint you're stuck on and you'll get a written scope and a fixed quote before any commitment.

Common questions

When was Grok 4.6 released?

SpaceXAI released Grok 4.6 on 12 August 2026, five weeks after Grok 4.5. It is available the same day in Cursor, Grok Build, the SpaceXAI API, and through partners including OpenRouter, Vercel and Cloudflare. Cursor and Grok Build are offering extra included usage for the first week.

How much does Grok 4.6 cost?

On the SpaceXAI API, prompts under 200,000 tokens cost $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. Once a prompt reaches 200,000 tokens, the whole request is billed at $4.00 / $1.00 / $12.00. A fast or priority tier is twice those rates. Cache hits are more expensive than on Grok 4.5, which billed them at $0.30 per million. Cursor's on-demand rates match the standard card, with a 50% launch discount for one week from 12 August 2026.

Is Grok 4.6 better than GPT-5.6 Sol?

On the Artificial Analysis Intelligence Index they are tied at 61. Grok 4.6 leads or matches Sol on GDPVal-AA, CursorBench, FrontierCode, APEX-Agents and AA-Briefcase in SpaceXAI's launch table, and trails on DeepSWE v1.1 (65.9% vs 73%) and Terminal-Bench v3.0 (26% vs 34.6%). Sol still costs $5/$30 per million tokens against Grok's $2/$6. The useful comparison is cost per completed task on your harness, not the composite index.

What is the difference between Grok 4.6 and Grok 4.5?

Grok 4.6 is a post-training upgrade on the same 1.5-trillion-parameter foundation, not a larger model. Headline token prices are unchanged at $2/$6, but cache hits rose from $0.30 to $0.50 per million. The Intelligence Index moved from 56 to 61. The largest published gains are on long agent and knowledge-work tests; DeepSWE and Terminal-Bench v3.0 improved but still trail GPT-5.6 Sol and Claude Fable 5. SpaceXAI also reports more self-checking on long runs and a stronger first pass on visual and interactive apps.

Does Grok 4.6 work in Cursor, and is it available in the EU?

Yes. Grok 4.6 is a first-party Cursor model, billed from the Cursor Models pool on paid plans, and it is available in the desktop app, cloud agents, CLI, SDK, Automations and the iOS app. Cursor's help docs say it is offered in every country where Cursor ships models, including the EU. Effort defaults to high; the Start plan (India) is fixed at medium effort without Fast mode. Enterprise admins who previously blocked the xAI provider can still enable 4.6 from the Cursor vendor catalogue.

What is Grok 4.6's context window?

500,000 tokens, the same as Grok 4.5. Prompts at or above 200,000 tokens are billed at double the standard rates for the entire request. SpaceXAI recommends setting a prompt_cache_key so multi-turn loops hit the same server, and using context compaction on long agent jobs so you do not wander into the expensive tier by accident.

Working on something like this?

Tell us what you're building. You'll get a written scope and a fixed quote before any commitment.

Get a fixed quote

Practical notes on shipping software

Occasional, concrete write-ups on building AI, web and mobile products: the kind of thing we'd tell a founder on a call. No spam, unsubscribe anytime.