AI/TLDR

Gemini 3.6 Flash

Google's July 2026 Flash workhorse — 17% fewer output tokens than 3.5 Flash with stronger agentic, coding and long-context scores.

Overview

Gemini 3.6 Flash is Google's Flash-tier workhorse model, released July 21, 2026 as the API model gemini-3.6-flash. It is a natively multimodal reasoning model that accepts text, images, audio, video, and PDFs as input with a 1 million-token context window and up to 64K output tokens, and it is Google's headline replacement for Gemini 3.5 Flash on scaled agentic and knowledge-work loads.

The signature win Google emphasises is efficiency: Gemini 3.6 Flash uses about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Intelligence Index and up to 65% fewer on DeepSWE, while still improving on the Flash predecessor across coding, computer use, long context, and multimodal benchmarks. Google's model card reports 58.7% on SWE-Bench Pro (vs 55.1% for 3.5 Flash and 54.2% for 3.1 Pro), 78% on Terminal-Bench 2.1, 83.0% on OSWorld-Verified, and 54.0% on GDM-MRCR v2 at the full 1M-token depth — roughly double the 3.5 Flash and 3.1 Pro scores at the same depth.

The model is available through the Gemini API, Google AI Studio, Vertex AI, the Gemini app, Google Search AI Mode, and Google's Antigravity agentic IDE. Its knowledge cutoff is March 2026; for newer information Google recommends the built-in search grounding tool. It is paired at launch with Gemini 3.5 Flash-Lite (the new cheapest Gemini 3 tier) and Gemini 3.5 Flash Cyber (a limited-access cybersecurity variant).

Released2026-07-21
LicenseProprietary
WeightsAPI only
ParametersNot disclosed
Context1M
Max output64K
ArchitectureNatively multimodal Gemini model in the Flash tier, positioned by Google as its workhorse for scaled agentic and knowledge work. Google reports 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index and up to 65% token reduction on DeepSWE, alongside built-in tool use (function calling, structured output, code execution, computer use, and search grounding).
Knowledge cutoffMarch 2026
ModalitiesText, Vision, Audio, Video, PDF
StatusGenerally available

Benchmarks

Gemini 3.6 Flash benchmark comparison (Google launch numbers, July 21, 2026).

BenchmarkGemini 3.6 FlashGemini 3.5 FlashGemini 3.1 Pro
SWE-Bench Pro (agentic coding)58.7%55.1%54.2%
Terminal-Bench 2.1 (agentic terminal coding)78%76.2%73.8%
DeepSWE v1.149%37%12%
MLE-Bench63.9%49.7%42.6%
OSWorld-Verified (agentic computer use)83%78.4%
GDPval-AA v2 (economically valuable knowledge work)1421 Elo1349 Elo965 Elo
GDM-MRCR v2 (1M-token long context)54%

Comparison source ↗

This model's scores

  1. SWE-Bench Pro58.7%
  2. Terminal-Bench 2.1 (agentic coding)78%
  3. DeepSWE v1.149%
  4. MLE-Bench63.9%
  5. OSWorld-Verified (agentic computer use)83%
  6. GDM-MRCR v2 (1M-token long context)54%
  7. GDPval-AA v2 (economically valuable knowledge work)1421Elo

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$1.50 / 1M tokens
Cached input$0.15 / 1M tokens
Output$7.50 / 1M tokens

Standard paid tier; output price includes thinking tokens. Cached input storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.

Pricing source ↗

Strengths

  • 17% lower output-token usage than Gemini 3.5 Flash on the Artificial Analysis Index (up to 65% fewer on DeepSWE) — the same or better answer at meaningfully lower cost
  • 58.7% on SWE-Bench Pro — beats both Gemini 3.5 Flash (55.1%) and Gemini 3.1 Pro (54.2%)
  • 83.0% on OSWorld-Verified and 78% on Terminal-Bench 2.1 for agentic computer and terminal work
  • 54.0% on GDM-MRCR v2 at the full 1M-token depth — roughly double 3.5 Flash and 3.1 Pro at the same depth
  • 1M-token multimodal context (text, image, audio, video, PDF) with 64K output
  • Built-in tool use: function calling, structured output, code execution, computer use, and search grounding

Best for

  • Reach for it as the default Flash-tier workhorse for agentic coding, computer use, and multi-step tool workflows
  • Reach for it when output-token cost dominates the bill — its lower token use compounds across a scaled agent loop
  • Reach for it for long-document and repo-scale reasoning that needs the full 1M-token window
  • Reach for it for multimodal tasks that mix images, audio, video, and PDFs with text

How to access

ProviderModel ID
Google Gemini API ↗gemini-3.6-flash
Google Vertex AI ↗gemini-3.6-flash

Gemini Flash — every version

The full lineage of the Gemini Flash line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Gemini 3.6 Flashcurrent2026-07-211MProprietary
Gemini 3.5 Flash2026-05-19Proprietary
Gemini 3 Flash2025-12-17Proprietary
Gemini 2.5 Flash2025-04-17Proprietary
Gemini 2.0 Flash2025-01-30Proprietary
Gemini 1.5 Flash2024-05-14Proprietary

FAQ

When was Gemini 3.6 Flash released?

Google released Gemini 3.6 Flash on July 21, 2026 as the API model gemini-3.6-flash, alongside Gemini 3.5 Flash-Lite and the limited-access Gemini 3.5 Flash Cyber. It is the successor to Gemini 3.5 Flash in the Flash tier.

What is the context window of Gemini 3.6 Flash?

Gemini 3.6 Flash has a 1 million-token input context window and can produce up to 64,000 output tokens. It accepts text, images, audio, video, and PDFs as input and returns text.

How much does Gemini 3.6 Flash cost?

On the standard paid tier of the Gemini API, Gemini 3.6 Flash is $1.50 per million input tokens and $7.50 per million output tokens, with cached input at $0.15 per million tokens. Context-cache storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.

How does Gemini 3.6 Flash compare to Gemini 3.5 Flash?

Google reports Gemini 3.6 Flash uses about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Intelligence Index — and up to 65% fewer on DeepSWE — while still improving on the older model across coding (SWE-Bench Pro 58.7% vs 55.1%), agentic terminal use (Terminal-Bench 2.1 78% vs 76.2%), computer use (OSWorld-Verified 83.0% vs 78.4%), and 1M-token long context (GDM-MRCR v2 54.0% vs under 27%).