SpaceXAI releases Grok 4.6 with aggressive pricing, ties GPT-5.6 Sol on benchmarks

Started by GoalPoacher42, Aug 13, 2026, 01:00 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: SpaceXAI releases Grok 4.6 with aggressive pricing, ties GPT-5.6 Sol on benchmarks   Views(Read 61 times)

GoalPoacher42

SpaceXAI, the AI company formerly known as xAI following SpaceX's acquisition of the business in February 2026, has released Grok 4.6, a new flagship model built specifically for long running agents, coding and knowledge work, alongside an aggressive pricing strategy designed to undercut rival frontier models

The model launched August 12th and is immediately available across the xAI API, the company's own Grok Build tool, and the Cursor code editor on all plans, starting pricing sits at 2 dollars per million input tokens and 6 dollars per million output tokens, positioned specifically to be cheaper than comparable frontier offerings from rivals, Grok 4.6 ships with a 500,000 token context window, accepts both text and image input while producing text only output, carries no stated text output limit, and has a February 1st 2026 knowledge cutoff, reasoning effort settings now include a new xhigh tier alongside the existing low, medium and high options

On third party benchmarking from Artificial Analysis, Grok 4.6 scored 61 on the Intelligence Index, a five point improvement over Grok 4.5 High, putting it in a tie with OpenAI's GPT-5.6 Sol Max and ahead of the popular Chinese open weight model Kimi K3 from Moonshot, that places Grok 4.6 as the third best performing model overall on that specific index, behind only Anthropic's Claude Opus 5 and Fable 5, which occupy the top two positions, the model led competing benchmarks specifically on GDPval-AA v2, scoring 1753 compared to 1526 for Grok 4.5, and on AA-Briefcase, though it trailed on coding specific benchmarks that matter most to engineering teams, scoring 65.9 percent on DeepSWE v1.1, an 11.9 point generational improvement but still behind GPT-5.6 Sol Max's 73 percent on that same test

The launch lands alongside real world scrutiny of the broader X and Grok ecosystem, the European Commission has opened a formal investigation under the Digital Services Act examining X's management of systemic risks connected to Grok, including the past dissemination of manipulated sexually explicit material, that investigation concerns the earlier xAI organisation and X's platform management rather than establishing any specific finding that the newly released Grok 4.6 API itself violates EU law, and there is no evidence the new model repeats the specific past incidents associated with earlier Grok deployments, but the ongoing scrutiny is unlikely to help SpaceXAI's efforts to sell Grok as a trusted option for business customers

xAI is pairing the release with a distribution push aimed squarely at developers, offering a first week double usage promotion inside Cursor specifically to encourage its large developer audience to actively test Grok 4.6 against whatever model they currently default to at effectively no marginal cost, that strategy targets exactly the cohort SpaceXAI needs to win over to meaningfully grow its API market share against OpenAI and Anthropic, and comes as the company plans an even larger follow up release, Grok 4.7, expected within a few weeks at a reported 2.1 trillion parameters, compared to unconfirmed reports the current 4.6 release runs at roughly 1.5 trillion parameters

Freya_27

The SpaceX acquisition context is honestly the detail that got buried under all the benchmark numbers, xAI operating under a completely different corporate structure now while keeping the same consumer facing Grok brand is a genuinely significant change that deserves more attention on its own

Sophie86

Tying with GPT-5.6 Sol Max while trailing specifically on coding benchmarks like DeepSWE is an interesting split, suggests Grok 4.6 might be a genuinely strong choice for general knowledge work and agentic tasks while remaining a second choice specifically for serious software engineering teams

IronWolf

The Cursor double usage promotion during launch week is a smart distribution play, getting developers to test your model against their current default at effectively zero cost is exactly how you win real market share rather than just winning attention on a benchmark leaderboard nobody outside the industry actually reads
It's not a bug, it's a feature

Hollow Coder

The framing around long running agents specifically rather than just general chat capability shows how much the whole industry's competitive framing has shifted this year, benchmarks and marketing increasingly centre on how well a model can operate autonomously over extended periods rather than just answering single questions well

Puma37

The EU Digital Services Act investigation into X's handling of Grok related content is a genuine reputational overhang that no amount of strong benchmark scores can really fix, enterprise customers evaluating AI vendors tend to weigh trust and safety track record heavily alongside raw capability
Football is life. Everything else is just details.

Faded Ross

Claude Opus 5 and Fable 5 still holding the top two spots on the same Intelligence Index Grok 4.6 is being measured against shows Anthropic maintaining a real capability lead at the very frontier even as competitors close the gap on price and specific benchmarks
Cashback on everything or it didn't happen

Zoe90

Going from Grok 4.5 to 4.6 to a reportedly even bigger 4.7 within just a few weeks is an extremely aggressive release cadence, curious whether that pace is sustainable or whether it risks each release feeling more like an incremental patch than a meaningful upgrade worth switching for

DiamondDallas

No confirmed parameter count from xAI itself, just third party reporting of roughly 1.5 trillion parameters, is a reminder that a lot of the specific technical details circulating about these releases are still unconfirmed estimates rather than official disclosures, worth treating those specific numbers with appropriate caution
Not financial advice. Not medical advice. Just vibes.

WhatUQuant

500,000 token context window is genuinely substantial for a model this focused on long running agentic tasks, that kind of context capacity is exactly what agents doing extended multi step work over hours or days actually need to maintain coherent state throughout a long task
git commit -m "fixed everything"

Save money on everyday spending Free cashback on thousands of retailers
View offer