Overview
Gemini 3.6 Flash is Google's Flash-tier workhorse model, released July 21, 2026 as the API model gemini-3.6-flash. It is a natively multimodal reasoning model that accepts text, images, audio, video, and PDFs as input with a 1 million-token context window and up to 64K output tokens, and it is Google's headline replacement for Gemini 3.5 Flash on scaled agentic and knowledge-work loads.
The signature win Google emphasises is efficiency: Gemini 3.6 Flash uses about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Intelligence Index and up to 65% fewer on DeepSWE, while still improving on the Flash predecessor across coding, computer use, long context, and multimodal benchmarks. Google's model card reports 58.7% on SWE-Bench Pro (vs 55.1% for 3.5 Flash and 54.2% for 3.1 Pro), 78% on Terminal-Bench 2.1, 83.0% on OSWorld-Verified, and 54.0% on GDM-MRCR v2 at the full 1M-token depth — roughly double the 3.5 Flash and 3.1 Pro scores at the same depth.
The model is available through the Gemini API, Google AI Studio, Vertex AI, the Gemini app, Google Search AI Mode, and Google's Antigravity agentic IDE. Its knowledge cutoff is March 2026; for newer information Google recommends the built-in search grounding tool. It is paired at launch with Gemini 3.5 Flash-Lite (the new cheapest Gemini 3 tier) and Gemini 3.5 Flash Cyber (a limited-access cybersecurity variant).
| Released | 2026-07-21 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Parameters | Not disclosed |
| Context | 1M |
| Max output | 64K |
| Architecture | Natively multimodal Gemini model in the Flash tier, positioned by Google as its workhorse for scaled agentic and knowledge work. Google reports 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index and up to 65% token reduction on DeepSWE, alongside built-in tool use (function calling, structured output, code execution, computer use, and search grounding). |
| Knowledge cutoff | March 2026 |
| Modalities | Text, Vision, Audio, Video, PDF |
| Status | Generally available |
Benchmarks
Gemini 3.6 Flash benchmark comparison (Google launch numbers, July 21, 2026).
| Benchmark | Gemini 3.6 Flash | Gemini 3.5 Flash | Gemini 3.1 Pro |
|---|---|---|---|
| SWE-Bench Pro (agentic coding) | 58.7% | 55.1% | 54.2% |
| Terminal-Bench 2.1 (agentic terminal coding) | 78% | 76.2% | 73.8% |
| DeepSWE v1.1 | 49% | 37% | 12% |
| MLE-Bench | 63.9% | 49.7% | 42.6% |
| OSWorld-Verified (agentic computer use) | 83% | 78.4% | — |
| GDPval-AA v2 (economically valuable knowledge work) | 1421 Elo | 1349 Elo | 965 Elo |
| GDM-MRCR v2 (1M-token long context) | 54% | — | — |
This model's scores
- SWE-Bench Pro58.7%
- Terminal-Bench 2.1 (agentic coding)78%
- DeepSWE v1.149%
- MLE-Bench63.9%
- OSWorld-Verified (agentic computer use)83%
- GDM-MRCR v2 (1M-token long context)54%
- GDPval-AA v2 (economically valuable knowledge work)1421Elo
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $1.50 / 1M tokens |
|---|---|
| Cached input | $0.15 / 1M tokens |
| Output | $7.50 / 1M tokens |
Standard paid tier; output price includes thinking tokens. Cached input storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.
Strengths
- 17% lower output-token usage than Gemini 3.5 Flash on the Artificial Analysis Index (up to 65% fewer on DeepSWE) — the same or better answer at meaningfully lower cost
- 58.7% on SWE-Bench Pro — beats both Gemini 3.5 Flash (55.1%) and Gemini 3.1 Pro (54.2%)
- 83.0% on OSWorld-Verified and 78% on Terminal-Bench 2.1 for agentic computer and terminal work
- 54.0% on GDM-MRCR v2 at the full 1M-token depth — roughly double 3.5 Flash and 3.1 Pro at the same depth
- 1M-token multimodal context (text, image, audio, video, PDF) with 64K output
- Built-in tool use: function calling, structured output, code execution, computer use, and search grounding
Best for
- Reach for it as the default Flash-tier workhorse for agentic coding, computer use, and multi-step tool workflows
- Reach for it when output-token cost dominates the bill — its lower token use compounds across a scaled agent loop
- Reach for it for long-document and repo-scale reasoning that needs the full 1M-token window
- Reach for it for multimodal tasks that mix images, audio, video, and PDFs with text
How to access
| Provider | Model ID |
|---|---|
| Google Gemini API ↗ | gemini-3.6-flash |
| Google Vertex AI ↗ | gemini-3.6-flash |
Gemini Flash — every version
The full lineage of the Gemini Flash line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Gemini 3.6 Flashcurrent | 2026-07-21 | 1M | Proprietary |
| Gemini 3.5 Flash | 2026-05-19 | — | Proprietary |
| Gemini 3 Flash | 2025-12-17 | — | Proprietary |
| Gemini 2.5 Flash | 2025-04-17 | — | Proprietary |
| Gemini 2.0 Flash | 2025-01-30 | — | Proprietary |
| Gemini 1.5 Flash | 2024-05-14 | — | Proprietary |
FAQ
When was Gemini 3.6 Flash released?
Google released Gemini 3.6 Flash on July 21, 2026 as the API model gemini-3.6-flash, alongside Gemini 3.5 Flash-Lite and the limited-access Gemini 3.5 Flash Cyber. It is the successor to Gemini 3.5 Flash in the Flash tier.
What is the context window of Gemini 3.6 Flash?
Gemini 3.6 Flash has a 1 million-token input context window and can produce up to 64,000 output tokens. It accepts text, images, audio, video, and PDFs as input and returns text.
How much does Gemini 3.6 Flash cost?
On the standard paid tier of the Gemini API, Gemini 3.6 Flash is $1.50 per million input tokens and $7.50 per million output tokens, with cached input at $0.15 per million tokens. Context-cache storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.
How does Gemini 3.6 Flash compare to Gemini 3.5 Flash?
Google reports Gemini 3.6 Flash uses about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Intelligence Index — and up to 65% fewer on DeepSWE — while still improving on the older model across coding (SWE-Bench Pro 58.7% vs 55.1%), agentic terminal use (Terminal-Bench 2.1 78% vs 76.2%), computer use (OSWorld-Verified 83.0% vs 78.4%), and 1M-token long context (GDM-MRCR v2 54.0% vs under 27%).