xAI · 2026-08-12 · seismic
Grok 4.6 — xAI's new flagship built for long-running agents
Grok 4.6 is xAI's new flagship, tuned for agents that run long multi-step tasks. It scores 65.9% on DeepSWE v1.1, up from Grok 4.5's 54%, and keeps the same $2/$6 per million tokens.

xAI's new flagship, improved through post-training rather than size, so it holds up over long agent runs.
Quick facts
| Maker | xAI |
|---|---|
| Context window | 500K tokens |
| Price (input) | $2 / 1M tokens |
| Price (output) | $6 / 1M tokens |
| Availability | API, Cursor, Grok Build, OpenRouter |
| Knowledge cutoff | February 1, 2026 |
| What's new | Post-training for long agent runs, same 1.5T foundation |
Benchmarks
Pricing
| Input · prompts under 200K tokens | $2.00 / 1M tokens |
|---|---|
| Cached input · prompts under 200K tokens | $0.50 / 1M tokens |
| Output · prompts under 200K tokens | $6.00 / 1M tokens |
| Input · prompts of 200K tokens or more | $4.00 / 1M tokens |
| Output · prompts of 200K tokens or more | $12.00 / 1M tokens |
What is it?
Grok 4.6 targets work that takes many steps — research, analysis, changes across a codebase, and building whole apps. xAI shipped it on 12 August 2026 into Cursor, Grok Build, the API console, and partners including OpenRouter, Vercel and Cloudflare. Price sits where Grok 4.5 sat, at $2 per million input tokens and $6 per million output tokens.
How does it work?
The gains come from post-training, not a bigger model. xAI ran a longer supplemental training pass over curated model-generated data covering reasoning and technical concepts, used an improved optimizer, and applied reinforcement learning on agentic tasks such as kernel optimization, web development and computer-aided design. xAI says the result is a model that self-tests and verifies its own work part-way through a long run.
Why does it matter?
Long agent runs are where models usually fall apart, and that is where the numbers move: DeepSWE v1.1 goes from 54% on Grok 4.5 to 65.9%, and Terminal-Bench v3.0 from 15.7% to 26%. At $2/$6 per million tokens, Grok 4.6 costs far less than Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30, while matching GPT-5.6 Sol's score of 61 on the Artificial Analysis Intelligence Index.
Who is it for?
developers running long coding and research agents
Frequently asked questions
- How much does Grok 4.6 cost?
- Grok 4.6 costs $2 per million input tokens and $6 per million output tokens for prompts under 200K tokens, with cached input at $0.50. Prompts of 200K tokens or more double to $4 in and $12 out. A faster variant costs double the standard rate. xAI is giving 2x included usage for the first week in Grok Build and Cursor.
- What changed between Grok 4.5 and Grok 4.6?
- Grok 4.6 keeps the same foundation model as Grok 4.5 and puts the work into post-training instead: a longer supplemental training run, an improved optimizer, and reinforcement learning on agentic tasks. The benchmark gaps show up on agent work — DeepSWE v1.1 rises from 54% to 65.9%, Terminal-Bench v3.0 from 15.7% to 26%, and APEX-Agents from 47.1% to 57.5%.
- Where can I use Grok 4.6 today?
- Grok 4.6 launched immediately in Cursor, in xAI's own Grok Build tool, and in the xAI API console. It is also served through partners including OpenRouter, Vercel and Cloudflare, so most existing routing setups can reach it without a direct xAI account.
- How does Grok 4.6 compare to GPT-5.6 Sol and Fable 5?
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max and just behind Fable 5 Max at 62. On coding agents it trails both — DeepSWE v1.1 puts Grok 4.6 at 65.9% against 73% for GPT-5.6 Sol Max. On Harvey LAB it leads clearly, at 15.8% versus 2.5% and 11.3%.
Try it
Pick Grok 4.6 in Cursor or Grok Build — 2x included usage for the first week