xAI releases Grok 4.6, ranking third globally on Artificial Analysis index
xAI, formerly known as SpaceXAI, has released Grok 4.6, scoring 61 on the Artificial Analysis Intelligence Index — matching OpenAI's GPT-5.6 Sol Max and surpassing Moonshot's Kimi K3 to share third place globally. The model improves five points over its predecessor Grok 4.5 High, with gains in coding, terminal, knowledge-work and agent benchmarks. API pricing starts at $2 per million input tokens and $6 per million output tokens. Anthropic's Claude Opus 5 and Fable 5 hold the top two spots.
Full text
Elon Musk's company SpaceXAI, formerly known as xAI, has released Grok 4.6 , its latest frontier AI model, with a focus on long-running agents, coding and knowledge work — and a pricing strategy designed to make those workloads cheaper to run.
The model scores 61 on the third-party Artificial Analysis Intelligence Index , surpassing the popular open weights Chinese model from Moonshot, Kimi K3, and tying rival OpenAI's GPT-5.6 Sol Max and improving five points over Grok 4.5 High. Anthropic's Claude Opus 5 and Fable 5 occupy the number one and two spots, respectively.
More consequential for enterprises evaluating AI agents, Grok 4.6 posts sizable gains over its predecessor across coding, terminal, knowledge-work and agent benchmarks while retaining an application programming interface (API) price starting at $2 per million input tokens and $6 per million output tokens, making it a mid-priced frontier model comparing leading options that are both proprietary and open source, globally, according to VentureBeat's analysis.
Model
Input ($/1M)
Output ($/1M)
Total ($/1M)
Source
Muse Spark 1.2 Contributor
$0.10
$0.20
$0.30
Meta
MiMo-V2.5 Flash
$0.10
$0.30
$0.40
Xiaomi
deepseek-v4-flash
$0.14
$0.28
$0.42
DeepSeek
deepseek-v4-pro
$0.435
$0.87
$1.305
DeepSeek
GPT-5.6 Luna
$0.20
$1.20
$1.40
OpenAI
MiniMax-M3
$0.30
$1.20
$1.50
MiniMax
LongCat-2.0 — limited-time promo
$0.30
$1.20
$1.50
LongCat
MiMo-V2.5
$0.40
$2.00
$2.40
Xiaomi
LongCat-2.0 — standard
$0.75
$2.95
$3.70
LongCat
MiMo-V2.5 Pro (≤256K)
$1.00
$3.00
$4.00
Xiaomi
Muse Spark 1.1 / 1.2
$1.25
$4.25
$5.50
Meta
GLM-5.2
$1.40
$4.40
$5.80
Z.ai
Grok 4.6 — <200K prompt tokens
$2.00
$6.00
$8.00
xAI
MiMo-V2.5 Pro (>256K)
$2.00
$6.00
$8.00
Xiaomi
Qwen3.8-Max
$2.00
$6.00
$8.00
QwenCloud
Gemini 3.6 Flash
$1.50
$7.50
$9.00
GPT-5.6 Terra
$2.00
$12.00
$14.00
OpenAI
Grok 4.6 — ≥200K prompt tokens
$4.00
$12.00
$16.00
xAI
GPT-5.4
$2.50
$15.00
$17.50
OpenAI
Kimi K3
$3.00
$15.00
$18.00
Moonshot AI
Claude Opus 5
$5.00
$25.00
$30.00
Anthropic
Sakana Fugu Ultra (≤272K)
$5.00
$30.00
$35.00
Sakana AI
GPT-5.6 Sol — Standard mode
$5.00
$30.00
$35.00
OpenAI
Claude Fable 5 / Claude Mythos 5
$10.00
$50.00
$60.00
Anthropic
GPT-5.6 Sol — Fast mode
$10.00
$60.00
$70.00
OpenAI
Still, that's less than half of what GPT-5.6 Sol costs over OpenAI's API in standard mode.
SpaceXAI says Grok 4.6 is available today in Grok Build, SpaceXAI's answer to Anthropic's Claude Code and OpenAI's Codex, which is available starting in the $30 per month SuperGrok plan .
It's also available in SpaceX's recent acquisition of the AI coding startup Cursor, and from partners including OpenRouter, Vercel and Cloudflare.
SpaceXAI is providing twice the included usage for Grok 4.6 in Cursor and Grok Build during the first week.
The release arrives only weeks after Grok 4.5, which SpaceXAI launched in July as a model targeting coding, agentic tasks and knowledge work, and one day after the launch of Grok Bot , a new system for assigning AI agents to complete designated tasks as virtual employees.
The bigger change is agent behavior, not just another benchmark point
SpaceXAI describes Grok 4.6 as being built specifically to stay on task across longer sequences of work, including researching unfamiliar topics, analyzing information, navigating codebases and converting product ideas into working applications.
The company says it subjected the model to a longer supplemental training run than Grok 4.5, using curated model-generated reasoning and technical data alongside engineering data and changes to its optimizer and training recipe. It then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning levels, agent harnesses, STEM, software engineering and knowledge work, filtering problematic trajectories with model-based checks.
Reinforcement learning also targeted agentic environments spanning general coding, knowledge work, kernel optimization, web development and computer-aided design.
That matters because enterprise AI deployments are increasingly moving beyond isolated prompt-and-response interactions toward agents expected to maintain state, operate tools, modify code and recover from problems across longer execution paths.
SpaceXAI says that during its testing, Grok 4.6 showed more self-testing and verification on longer trajectories, checking its own work before proceeding. It also reports stronger first attempts on interactive and visual projects than Grok 4.5. Those are company observations rather than independent guarantees of production behavior, but they indicate where SpaceXAI concentrated the model’s post-training work.
Grok 4.6 reaches the frontier, but does not sweep it
Grok 4.6's improvement over Grok 4.5 at this juncture of the AI model competition cannot be overstated.
According to Artificial Analysis, Grok 4.6 reaches an Elo score (human preference of head-to-head model outputs, adapted from chess) of 1,753 on GDPVal-AA v2 , the benchmark measuring performance on real-world tasks like scheduling and diagramming, versus 1,526 for Grok 4.5, 1,728 for GPT-5.6 Sol Max and 1,741 for Fable 5 Max.
The coding results from SpaceXAI show a similar generational improvement but more competition at the frontier.
Grok 4.6 scores 69.9% on CursorBench v3.2, up from 66.7%, while Fable 5 Max reaches 70.5%. On DeepSWE v1.1, Grok rises sharply from 54% to 65.9%, but GPT-5.6 Sol Max leads at 73%. FrontierCode v1.1 Extended moves from 56.6% to 61.3%, compared with 60.6% for GPT-5.6 Sol Max and a leading 63.6% for Fable 5 Max.
Agent benchmarks tell much the same story. Grok 4.6 reaches 57.5% on APEX-Agents, a 10.4-point increase over Grok 4.5’s 47.1%, narrowly exceeding GPT-5.6 Sol Max’s 56.7% but trailing Fable 5 Max at 59.2%. On APEX-SWE, Grok 4.6 rises to 56.4% from 53.6%, while Fable 5 Max scores 58.8%.
Terminal-Bench v3.0 exposes a larger remaining gap. Grok 4.6 improves from 15.7% to 26%, but GPT-5.6 Sol Max and Fable 5 Max score 34.6% and 34.1%, respectively.
Two of Grok 4.6’s strongest results come from longer-horizon professional work. On AA-Briefcase it scores an Elo of 1,577, narrowly exceeding Fable 5 Max’s 1,574 and topping GPT-5.6 Sol Max’s 1,502. On Harvey LAB, Grok 4.6 reaches 15.8%, versus 12.9% for Grok 4.5, 11.3% for Fable 5 Max and 2.5% for GPT-5.6 Sol Max.
SpaceXAI notes an important methodological caveat: third-party scores in its table use the best self-reported or publicly available results. The comparison therefore should not be interpreted as a perfectly controlled four-model evaluation.
In other words, the evidence supports a substantial upgrade over Grok 4.5 more clearly than it supports across-the-board superiority over rival frontier models. Grok 4.6 wins several of the displayed evaluations while GPT-5.6 Sol Max and Fable 5 Max retain meaningful leads elsewhere.
Cost could be the more important enterprise benchmark
Artificial Analysis’ supplied evaluation adds another dimension: how much work the model performs for the money spent.
The testing places Grok 4.6 on its Intelligence-versus-Cost-per-Task Pareto frontier at a reported $0.84 per task — which actually makes it less of a bargain than its predecessor, Grok 4.5, and less economical than OpenAI's GPT-5.6 Luna, z.ai's GLM-5.2, and Meta's new Muse Spark 1.2, among other models.
Artificial Analysis also reports that Grok 4.6 completed its AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens on average, versus approximately 103 turns and 2 billion input tokens for Claude Opus 5 Max.
Those measurements do not prove that every production agent will use fewer tokens or finish twice as quickly. Agent costs depend heavily on harness design, prompts, tool calls, caching, retry behavior and the task itself. But they point toward an increasingly important enterprise metric: the cost of completing a workflow, rather than simply the cost of generating one million tokens. That distinction is central to SpaceXAI’s positioning.
The standard Grok 4.6 API starts at $2 per million input tokens and $6 per million output tokens, and SpaceXAI also offers a faster variant at twice the price.
The supplied API documentation adds an important caveat for long-context deployments. Grok 4.6 supports a 500,000-token context window, but prompts below 200,000 tokens are billed at $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens. Once a prompt reaches 200,000 tokens, those rates rise to $4, $1 and $12 respectively, with the higher pricing applying to all tokens in that request.
That means enterprises should not extrapolate the $2/$6 headline pricing across the model’s entire context window when estimating total cost of ownership.
Artificial Analysis says the standard headline rates remain more than 60% below the competing frontier-model prices it cites for Claude Opus 5 and GPT-5.6 Sol. The practical savings will depend on how many tokens each model consumes to complete the same workload.
The Grok name carries considerable baggage and controversy, separate from the general AI skepticism
Performance and price may not be the only hurdles SpaceXAI faces in converting Grok 4.6's benchmark gains into enterprise adoption. The Grok brand arrives with an unusually visible history of safety and governance controversies — including extremist and antisemitic outputs, politically skewed responses, exaggerated praise of Elon Musk and, more recently, the use of Grok's image-generation capabilities to produce non-consensual sexualized imagery. For companies with strict compliance, brand-safety or responsible-AI requirements, that history could become a procurement consideration separate from the technical capabilities of Grok 4.6 itself.
The most notorious text-generation episode came in July 2025, when Grok produced antisemitic posts, praised Adolf Hitler and in some responses referred to itself as " MechaHitler ." SpaceXAI's predecessor xAI subsequently said it was removing inappropriate posts and taking steps to prevent hate speech from being published by Grok.
Also in summer 2025, Grok began inserting references to an alleged "white genocide" in South Africa into answers to unrelated questions. xAI said an unauthorized modification to Grok's response software had directed the system to produce a particular response on a political topic while bypassing its normal review process. The company said the change violated its policies and subsequently pledged to publish Grok's system prompts and establish round-the-clock monitoring for problematic responses. The South African government has rejected claims that a genocide against white South Africans is taking place.
Grok's objectivity came under scrutiny again in November of the same year after the chatbot repeatedly produced implausibly flattering assessments of Musk . Among the examples reported at the time were claims placing Musk above elite athletes and historic intellectual figures. Musk said Grok had been manipulated through adversarial prompting into making "absurdly positive" statements about him. Whatever the underlying cause, the incident illustrated the reputational problem for an enterprise model whose outputs can become entangled with the public persona of the executive most closely associated with its developer.
The most serious controversy has involved image generation. In January 2026, U.K. regulator Ofcom opened a formal investigation into X after reports that the Grok account was being used to create and distribute undressed images of people and sexualized images of children. Ofcom said the material under examination could amount to non-consensual intimate-image abuse, pornography and child sexual abuse material. X subsequently said it had implemented measures intended to stop the Grok account from being used to create intimate images of people, but Ofcom said its investigation remained open.
The scrutiny extends beyond Ofcom. Britain's Information Commissioner's Office is investigating X and xAI over both the development and deployment of Grok, including whether personal data was handled lawfully and whether adequate safeguards existed to prevent harmful manipulated imagery. The European Commission, meanwhile, opened a separate formal investigation under the Digital Services Act examining X's management of systemic risks connected to Grok , including the dissemination of manipulated sexually explicit material.
Those investigations concern X and the earlier xAI organization rather than establishing a finding that the newly released Grok 4.6 API violates those laws. Nevertheless, they are unlikely to help SpaceXAI sell Grok to businesses.
SpaceX acquired xAI in February 2026, and the AI operation now markets itself as SpaceXAI, meaning Grok's newest models sit under a different corporate structure but retain the same consumer-facing brand.
There is no evidence in the material examined here that Grok 4.6 itself repeats the specific "MechaHitler," "white genocide," sexual-image or Musk-flattery incidents associated with earlier Grok deployments. But enterprise procurement teams rarely evaluate a model in isolation from its vendor and product history. For SpaceXAI, that means Grok 4.6 may have to demonstrate not only that it is cheaper or more capable than competing frontier models, but that the controls around it are sufficiently predictable for organizations that cannot afford their AI supplier to become a brand-safety event.
That continuity creates a potential adoption problem that benchmark tables cannot measure. Developers choosing a model for an internal coding agent may care primarily about price, latency and task completion. A bank, government agency, healthcare provider or consumer brand deploying the same model into customer-facing or regulated workflows may also have to consider vendor governance, content-safety controls, auditability and reputational exposure.
A model designed to be deployed, not just chatted with
Grok 4.6 supports text and image inputs with text output, function calling, structured outputs and reasoning, according to the supplied API specifications. Those specifications also list rate limits of 150 requests per second and 50 million tokens per minute, with API availability in us-east-1 and us-west-2 .
Cursor’s launch announcement similarly characterizes Grok 4.6 as designed for long-running agents and ambitious interactive and visual work, giving developers immediate access to the model inside an established coding-agent environment rather than requiring them to build a new harness around the API first.
For enterprise buyers, that distribution may matter almost as much as another leaderboard result. Models increasingly compete not just on reasoning scores but on whether developers can place them inside existing coding, research and operational workflows without destabilizing those workflows or dramatically increasing inference costs, as well as incurring any blowback from associating with a controversial brand.
Grok 4.6 does not establish an uncontested performance lead. Its launch instead presents a different proposition: frontier-level intelligence, large improvements over the previous generation, stronger long-running agent behavior and relatively aggressive token economics.
The next test will be whether the efficiency Artificial Analysis observes on controlled agentic workloads carries into production. If Grok 4.6 can consistently complete long-running coding and knowledge-work tasks with fewer turns and fewer tokens, the model’s most important benchmark may ultimately be the enterprise inference bill rather than the leaderboard.
Similar stories
💻 Technology
xAI launches Grok Bot — an always-on AI agent that works like an autonomous teammate
The Verge · 7h ago
💻 Technology
Musk: SpaceX AI revenue to surpass all other divisions by September, $500bn projected for 2027
Space.com · 20h ago
💻 Technology
Musk's xAI sues Grok users over CSAM and fights Minnesota nudification ban
Ars Technica · 14d ago
Similar stories
💻 Technology
xAI launches Grok Bot — an always-on AI agent that works like an autonomous teammate
The Verge · 7h ago
💻 Technology
Musk: SpaceX AI revenue to surpass all other divisions by September, $500bn projected for 2027
Space.com · 20h ago
💻 Technology
Musk's xAI sues Grok users over CSAM and fights Minnesota nudification ban
Ars Technica · 14d ago
Does Grok 4.6 pose a real threat to OpenAI's dominance in the AI market?
Comments
No comments yet
Comments
No comments yet — be the first to weigh in 👇
No comments yet. Be the first!