AI/TLDR

xAI · 2026-08-12 · seismic

Grok 4.6 — xAI's new flagship built for long-running agents

Grok 4.6 is xAI's new flagship, tuned for agents that run long multi-step tasks. It scores 65.9% on DeepSWE v1.1, up from Grok 4.5's 54%, and keeps the same $2/$6 per million tokens.

xAI announcement card for the Grok 4.6 model release

xAI's new flagship, improved through post-training rather than size, so it holds up over long agent runs.

Quick facts

MakerxAI
Context window500K tokens
Price (input)$2 / 1M tokens
Price (output)$6 / 1M tokens
AvailabilityAPI, Cursor, Grok Build, OpenRouter
Knowledge cutoffFebruary 1, 2026
What's newPost-training for long agent runs, same 1.5T foundation

Benchmarks

DeepSWE v1.1
Grok 4.665.9%
Grok 4.554%
GPT-5.6 Sol Max73%
Fable 5 Max70%
source ↗
Terminal-Bench v3.0
Grok 4.626%
Grok 4.515.7%
GPT-5.6 Sol Max34.6%
Fable 5 Max34.1%
source ↗
CursorBench v3.2
Grok 4.669.9%
Grok 4.566.7%
GPT-5.6 Sol Max67.2%
Fable 5 Max70.5%
source ↗
APEX-Agents
Grok 4.657.5%
Grok 4.547.1%
GPT-5.6 Sol Max56.7%
Fable 5 Max59.2%
source ↗

Pricing

Input · prompts under 200K tokens$2.00 / 1M tokens
Cached input · prompts under 200K tokens$0.50 / 1M tokens
Output · prompts under 200K tokens$6.00 / 1M tokens
Input · prompts of 200K tokens or more$4.00 / 1M tokens
Output · prompts of 200K tokens or more$12.00 / 1M tokens
source ↗

What is it?

Grok 4.6 targets work that takes many steps — research, analysis, changes across a codebase, and building whole apps. xAI shipped it on 12 August 2026 into Cursor, Grok Build, the API console, and partners including OpenRouter, Vercel and Cloudflare. Price sits where Grok 4.5 sat, at $2 per million input tokens and $6 per million output tokens.

How does it work?

The gains come from post-training, not a bigger model. xAI ran a longer supplemental training pass over curated model-generated data covering reasoning and technical concepts, used an improved optimizer, and applied reinforcement learning on agentic tasks such as kernel optimization, web development and computer-aided design. xAI says the result is a model that self-tests and verifies its own work part-way through a long run.

Why does it matter?

Long agent runs are where models usually fall apart, and that is where the numbers move: DeepSWE v1.1 goes from 54% on Grok 4.5 to 65.9%, and Terminal-Bench v3.0 from 15.7% to 26%. At $2/$6 per million tokens, Grok 4.6 costs far less than Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30, while matching GPT-5.6 Sol's score of 61 on the Artificial Analysis Intelligence Index.

Who is it for?

developers running long coding and research agents

Frequently asked questions

How much does Grok 4.6 cost?
Grok 4.6 costs $2 per million input tokens and $6 per million output tokens for prompts under 200K tokens, with cached input at $0.50. Prompts of 200K tokens or more double to $4 in and $12 out. A faster variant costs double the standard rate. xAI is giving 2x included usage for the first week in Grok Build and Cursor.
What changed between Grok 4.5 and Grok 4.6?
Grok 4.6 keeps the same foundation model as Grok 4.5 and puts the work into post-training instead: a longer supplemental training run, an improved optimizer, and reinforcement learning on agentic tasks. The benchmark gaps show up on agent work — DeepSWE v1.1 rises from 54% to 65.9%, Terminal-Bench v3.0 from 15.7% to 26%, and APEX-Agents from 47.1% to 57.5%.
Where can I use Grok 4.6 today?
Grok 4.6 launched immediately in Cursor, in xAI's own Grok Build tool, and in the xAI API console. It is also served through partners including OpenRouter, Vercel and Cloudflare, so most existing routing setups can reach it without a direct xAI account.
How does Grok 4.6 compare to GPT-5.6 Sol and Fable 5?
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max and just behind Fable 5 Max at 62. On coding agents it trails both — DeepSWE v1.1 puts Grok 4.6 at 65.9% against 73% for GPT-5.6 Sol Max. On Harvey LAB it leads clearly, at 15.8% versus 2.5% and 11.3%.

Try it

Pick Grok 4.6 in Cursor or Grok Build — 2x included usage for the first week

Sources · 4 outlets

Tags

  • grok-4-6
  • xai
  • llm
  • frontier-model
  • agents
  • coding
  • reasoning
  • api
  • cursor
  • benchmark

← All releases · Learn AI