Overview
Grok 4.6 is SpaceXAI's Grok flagship, announced on August 12, 2026. SpaceXAI describes it as building on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work: staying with a complex task across many steps, whether that is researching a topic, analysing information, working across a codebase, or turning an idea into a polished application or work artifact. On longer trajectories the company reports seeing more self-testing and verification, with the model checking its own work before moving on.
The model is served with a 500,000-token context window, accepts text and image input, and has a knowledge cutoff of February 1, 2026. Reasoning effort is configurable as low, medium, high (the default), or extra-high, and the model supports function calling, web search, X search, and code execution tools. Its API model id is grok-4.6.
SpaceXAI says Grok 4.6 went through a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe, then refined the result with reinforcement learning across agentic tasks spanning knowledge work, general programming, and domain-specific environments such as kernel optimization, web development, and computer-aided design. It is available in Grok Build, Cursor, the SpaceXAI API, and partner platforms including OpenRouter, Vercel, and Cloudflare.
| Released | 2026-08-12 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Context | 500K |
| Knowledge cutoff | February 1, 2026 |
| Modalities | Text, Vision |
| Status | Available |
Benchmarks
Grok 4.6 against Grok 4.5 and rival frontier models, as published by SpaceXAI at launch.
| Benchmark | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
This model's scores
- CursorBench v3.269.9%
- DeepSWE v1.165.9%
- FrontierCode v1.1 (Extended)61.3%
- APEX-Agents57.5%
- APEX-SWE56.4%
- Terminal-Bench v3.026%
- Harvey LAB (Vals)15.8%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $2.00 / 1M tokens |
|---|---|
| Cached input | $0.50 / 1M tokens |
| Output | $6.00 / 1M tokens |
Rates apply to prompts under 200K tokens; prompts of 200K tokens or more are billed at $4.00 input, $1.00 cached input, and $12.00 output per 1M tokens.
Strengths
- Scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol Max on SpaceXAI's published comparison
- Large agentic-coding gains over Grok 4.5 — DeepSWE v1.1 65.9% vs 54%, APEX-Agents 57.5% vs 47.1%
- 500,000-token context window with text and image input
- Reasoning effort dial from low to extra-high, plus function calling, web search, X search, and code execution
- $2.00 / $6.00 per 1M input/output tokens for prompts under 200K tokens
Best for
- Reach for it for long-running coding agents inside Cursor and Grok Build that must hold a task across many steps.
- Reach for it for turning a broad product idea into a working first version and refining it over several rounds of feedback.
- Reach for it for multi-step research in unfamiliar domains where the model has to structure its own work.
- Reach for it for tool-heavy agent workflows that combine function calling, web or X search, and code execution.
How to access
| Provider | Model ID |
|---|---|
| SpaceXAI API ↗ | grok-4.6 |
| Cursor ↗ | grok-4.6 |
Grok (flagship) — every version
The full lineage of the Grok (flagship) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Grok 4.6current | 2026-08-12 | 500K | Proprietary |
| Grok 4.5 | 2026-07-09 | — | Proprietary |
| Grok 4.3 | 2026-04-30 | 1M | Proprietary |
| Grok 4.20 | 2026-03 | — | Proprietary |
| Grok 4.1 | 2025-11-17 | — | Proprietary |
| Grok 4 | 2025-07-09 | — | Proprietary |
| Grok 3 | 2025-02-17 | — | Proprietary |
| Grok 2 | 2024-08-20 | — | Open weights |
| Grok 1.5 | 2024-05-15 | — | Proprietary |
| Grok 1 | 2023-11-03 | — | Apache-2.0 |
FAQ
When was Grok 4.6 released?
SpaceXAI announced Grok 4.6 on August 12, 2026, and made it available the same day in Cursor and Grok Build, in the SpaceXAI API, and through partners including OpenRouter, Vercel, and Cloudflare.
How much does Grok 4.6 cost?
For prompts under 200K tokens, SpaceXAI prices Grok 4.6 at $2.00 per 1M input tokens, $0.50 per 1M cached input tokens, and $6.00 per 1M output tokens. Prompts of 200K tokens or more are billed at double those rates: $4.00 input, $1.00 cached input, and $12.00 output per 1M tokens.
What is the Grok 4.6 context window?
Grok 4.6 has a 500,000-token context window and accepts text and image input. Its knowledge cutoff is February 1, 2026.
How does Grok 4.6 compare with Grok 4.5?
On SpaceXAI's published comparison, Grok 4.6 High scores 65.9% on DeepSWE v1.1 against 54% for Grok 4.5 High, 57.5% vs 47.1% on APEX-Agents, 69.9% vs 66.7% on CursorBench v3.2, and 26% vs 15.7% on Terminal-Bench v3.0. Its Artificial Analysis Intelligence Index score rises from 56 to 61.
What is Grok 4.6 best at?
SpaceXAI positions it for long-running agents and interactive, visual work — holding a complex task across many steps, turning a broad product idea into a working first version, and researching unfamiliar domains. The company reports more self-testing and verification on long trajectories.