Overview
Gemini 3.1 Deep Think is Google's specialized reasoning mode, dated February 2026 on its own model page. Google describes it as built on top of Gemini 3.1 Pro: rather than a separate base model, it is the configuration of that model that spends far more inference-time compute working a problem before it answers. Google positions it for science, research and engineering — the cases where being right matters more than being fast.
The published comparison puts it well ahead of the field on abstract and scientific reasoning. Against Gemini 3 Pro Preview Thinking (High), Opus 4.6 Thinking (Max) and GPT-5.2 Thinking (xhigh), it scores 84.6% on ARC-AGI-2 (ARC Prize verified) versus 31.1%, 68.8% and 52.9%; 48.4% on Humanity's Last Exam with no tools and 53.4% with search and code execution; and 81.5% on MMMU-Pro. On competition science it reports 81.5% on IMO 2025 mathematics, 87.7% on IPO 2025 physics theory, 82.8% on ICO 2025 chemistry theory, and 50.5% on the CMT-Benchmark for condensed-matter theory. Its Codeforces Elo without tools is 3455.
Google illustrates the mode with applied work rather than chat: optimising problems in materials science, catching errors in mathematical arguments, and prototyping in mechanical engineering. Access is through the Gemini app at gemini.google.com. Google's model page does not state a context window, knowledge cutoff or per-token pricing for this mode, so none is listed here.
| Released | 2026-02 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Modalities | Text, Vision |
| Status | Available in the Gemini app |
Benchmarks
Gemini 3.1 Deep Think vs the peers Google published alongside it
| Benchmark | Gemini 3.1 Deep Think | Gemini 3 Pro Preview Thinking (High) | Opus 4.6 Thinking (Max) | GPT-5.2 Thinking (xhigh) |
|---|---|---|---|---|
| ARC-AGI-2 (ARC Prize verified) | 84.6 | 31.1 | 68.8 | 52.9 |
| Humanity's Last Exam (text + MM, no tools) | 48.4 | 37.5 | 40 | 34.5 |
| Humanity's Last Exam (search + code execution) | 53.4 | 45.8 | 53.1 | 45.5 |
| MMMU-Pro (no tools) | 81.5 | 81 | 73.9 | 79.5 |
| IMO 2025 (mathematics) | 81.5 | 14.3 | — | 71.4 |
| Codeforces (no tools) | 3455 Elo | 2512 Elo | 2352 Elo | — |
| IPO 2025 (theory, physics) | 87.7 | 76.3 | 71.6 | 70.5 |
| CMT-Benchmark (condensed matter theory) | 50.5 | 39.5 | 17.1 | 41 |
| ICO 2025 (theory, chemistry) | 82.8 | 69.6 | — | 72 |
This model's scores
- ARC-AGI-2 (ARC Prize verified)84.6%
- IPO 2025 (theory, physics)87.7%
- ICO 2025 (theory, chemistry)82.8%
- IMO 2025 (mathematics)81.5%
- MMMU-Pro (no tools)81.5%
- Humanity's Last Exam (search + code execution)53.4%
- CMT-Benchmark (condensed matter theory)50.5%
- Humanity's Last Exam (text + MM, no tools)48.4%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- 84.6% on ARC-AGI-2 (ARC Prize verified) — more than double the Gemini 3 Pro Preview Thinking (High) result of 31.1%
- 48.4% on Humanity's Last Exam with no tools, rising to 53.4% with search and code execution
- Competition-grade science: 81.5% on IMO 2025 mathematics, 87.7% on IPO 2025 physics theory, 82.8% on ICO 2025 chemistry theory
- 3455 Codeforces Elo with no tools, ahead of every peer in Google's published comparison
- 50.5% on the CMT-Benchmark for condensed-matter theory, the strongest score in the same table
- Built on Gemini 3.1 Pro, so it inherits that model's multimodal understanding (81.5% MMMU-Pro)
Best for
- Research-level mathematics, including checking a proof or argument for errors
- Olympiad-grade physics and chemistry problems
- Materials-science and other scientific optimisation problems
- Mechanical-engineering prototyping and design reasoning
- Abstract reasoning tasks of the ARC-AGI-2 kind that defeat standard chat models
- Any task where extra inference-time compute is worth a slower answer
Gemini Deep Think — every version
The full lineage of the Gemini Deep Think line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Gemini 3.1 Deep Thinkcurrent | 2026-02 | — | Proprietary |
| Gemini 3 Deep Think | 2025-12-03 | — | Proprietary |
| Gemini 2.5 Deep Think | 2025-08-01 | — | Proprietary |
FAQ
Is Gemini 3.1 Deep Think a separate model?
No. Google describes it as its most specialized reasoning mode, built on top of Gemini 3.1 Pro. It is the configuration of that model that spends much more compute on a problem before answering.
How does it compare with Gemini 3 Pro on abstract reasoning?
In Google's published table it scores 84.6% on ARC-AGI-2 (ARC Prize verified) against 31.1% for Gemini 3 Pro Preview Thinking (High).
What are its competition-science results?
Google reports 81.5% on IMO 2025 mathematics, 87.7% on IPO 2025 physics theory and 82.8% on ICO 2025 chemistry theory.
Where can I use it?
Through the Gemini app at gemini.google.com. Google's model page does not publish separate per-token pricing for the mode.
What is its context window?
Google's Deep Think model page does not state a context window, knowledge cutoff or maximum output for this mode, so none is listed here.