Overview
Gemini 3.7 Flash is the Flash-tier Gemini model Google released on 13 August 2026 as the API model `gemini-3.7-flash`. Google positions it for complex coding, agentic workflows and reliable multi-step execution, and its model card describes it as built on Gemini 3.6 Flash with algorithmic improvements to the core reasoning foundation. It accepts text, image, audio and video input across a 1 million-token context window and returns up to 64K output tokens.
The gains Google published are concentrated in software engineering and agent work: 43.6% on FrontierCode 1.1 Main (up from 34.4% for Gemini 3.6 Flash), 65.3% on DeepSWE v1.1 (up from 48.6%), 85.8% on Terminal-bench 2.1 (up from 78.0%), and 30.4% on the private AutomationBench set (up from 17.0%). Google also reports 1588 Elo on Code Arena for web development, 97.0% on GDM-MRCR v2 (8-needle, 128k average) for long-context retrieval, and 34.0% on GDP.pdf for expert PDF comprehension.
Reasoning depth is selectable through three thinking levels — low for latency-sensitive work, medium as the default, and high for the hardest problems. The model is served through the Gemini API in Google AI Studio, Vertex-adjacent enterprise surfaces (Gemini Enterprise app and Agent Platform), Google Antigravity, Android Studio, and the Gemini app's Spark experience for AI Pro and Ultra subscribers. Its knowledge cutoff is March 2026, with some domains limited to January 2025.
Pricing is $0.75 per million input tokens and $3.75 per million output tokens (thinking tokens included) through 31 December 2026, rising to $1.50 / $7.50 on 1 January 2027. Context caching is $0.075 per million tokens through the end of 2026 ($0.15 after), with cache storage at $0.50 per million tokens per hour ($1.00 after).
| Released | 2026-08-13 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Parameters | Not disclosed |
| Context | 1M |
| Max output | 64K |
| Architecture | Google describes Gemini 3.7 Flash as the next iteration in the Gemini 3 family, built on Gemini 3.6 Flash with algorithmic improvements to its core reasoning foundation rather than a new pretraining run. It exposes three thinking levels — low, medium (default) and high — and carries the same suite of built-in tools as 3.6 Flash. |
| Knowledge cutoff | March 2026 |
| Modalities | Text, Vision, Audio, Video |
| Status | Generally available |
Benchmarks

Gemini 3.7 Flash against the field, as published by Google at launch (13 August 2026).
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | Claude Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |
|---|---|---|---|---|---|
| Input price ($/1M tokens) | $0.75* | $0.75* | $2.00 | $2.00 | $1.25 |
| Output price ($/1M tokens) | $3.75* | $3.75* | $10.00 | $12.00 | $4.25 |
| Artificial Analysis Intelligence Index | 56 | 52 | 55 | 57 | 57 |
| FrontierCode 1.1 Main (score) | 43.6% | 34.4% | 42.7% | 41.3% | — |
| DeepSWE v1.1 | 65.3% | 48.6% | 53.8% | 69.6% | 54.9% |
| Code Arena (Elo) | 1588 | 1538 | 1541 | 1523 | 1535 |
| Terminal-bench 2.1 | 85.8% | 78% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 | 14.9% | 5.4% | 14.6% | 20.8% | — |
| AutomationBench (private set) | 30.4% | 17% | 10.7% | 23.6% | — |
| GDPVal-AA v2 (Elo) | 1525 | 1422 | 1598 | 1578 | 1628 |
| Harvey LAB-AA | 90.7% | 85.1% | 90.1% | 85.2% | — |
| GDP.pdf | 34% | 22% | 28% | 24.7% | 16% |
| CharXiv Reasoning (no tools) | 84.5% | 85.2% | 77% | 85.9% | — |
| CharXiv Reasoning (with tools) | 88.7% | 89.4% | 88.3% | — | — |
| LVBench (long video understanding) | 85.4% | 84.2% | 68.5% | 78.9% | — |
| GDM-MRCR v2, 8-needle (128k average) | 97% | 91.8% | 81.5% | 93.5% | — |
| OSWorld-2.0 | 47.9% | 33.8% | — | 50.2% | — |
| Agent's Last Exam (pass rate) | 26.3% | 24.2% | 33.3% | 28% | — |
| HLE-Verified | 53.6% | 51.2% | 31% | 51.1% | — |
| BioMysteryBench (human solvable) | 87.1% | 80.6% | 87.5% | 83.8% | — |
| BioMysteryBench (human difficult) | 43.5% | 41.2% | 34.1% | 49.4% | — |
| LABBench2 (biology research tasks) | 82.1% | 76.1% | 80.1% | 81.2% | — |
This model's scores
- FrontierCode 1.1 Main (production code quality)43.6%
- DeepSWE v1.1 (long-horizon software engineering)65.3%
- Terminal-bench 2.1 (agentic terminal coding)85.8%
- AutomationBench (private set)30.4%
- OSWorld-2.0 (agentic computer use)47.9%
- GDM-MRCR v2, 8-needle (128k average)97%
- HLE-Verified (multidisciplinary expert reasoning)53.6%
- Harvey LAB-AA (complex legal workflows)90.7%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.75 / 1M tokens |
|---|---|
| Cached input | $0.075 / 1M tokens |
| Output | $3.75 / 1M tokens |
Introductory rate through 31 December 2026. From 1 January 2027: $1.50 input, $7.50 output, $0.15 context caching. Cache storage is $0.50 per 1M tokens per hour through 2026, $1.00 after.
Strengths
- 65.3% on DeepSWE v1.1 for long-horizon software engineering, against 48.6% for Gemini 3.6 Flash
- 85.8% on Terminal-bench 2.1 and 47.9% on OSWorld-2.0 for agentic terminal and computer use
- 97.0% on GDM-MRCR v2 (8-needle, 128k average) — the strongest long-context retrieval number in Google's own launch comparison
- 90.7% on Harvey LAB-AA for complex legal workflows and 34.0% on GDP.pdf for expert PDF comprehension
- 1M-token multimodal context (text, image, audio, video) with 64K output tokens
- Selectable thinking levels (low / medium / high) to trade latency against reasoning depth
- $0.75 / $3.75 per million tokens through 2026 — the same introductory rate as Gemini 3.6 Flash
Best for
- Reach for it for agentic coding loops that run long tool-using tasks to completion
- Reach for it for web and desktop app generation, where Google reports higher-fidelity code and better design adherence
- Reach for it for repo-scale and long-document reasoning that needs the full 1M-token window
- Reach for it for enterprise workflow automation, where its private AutomationBench score nearly doubles 3.6 Flash's
How to access
| Provider | Model ID |
|---|---|
| Gemini API (Google AI Studio) ↗ | gemini-3.7-flash |
Gemini Flash — every version
The full lineage of the Gemini Flash line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Gemini 3.7 Flashcurrent | 2026-08-13 | 1M | Proprietary |
| Gemini 3.6 Flash | 2026-07-21 | 1M | Proprietary |
| Gemini 3.5 Flash | 2026-05-19 | — | Proprietary |
| Gemini 3 Flash | 2025-12-17 | — | Proprietary |
| Gemini 2.5 Flash | 2025-04-17 | — | Proprietary |
| Gemini 2.0 Flash | 2025-01-30 | — | Proprietary |
| Gemini 1.5 Flash | 2024-05-14 | — | Proprietary |
FAQ
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is a Flash-tier Gemini model Google released on 13 August 2026 as the API model gemini-3.7-flash. Google's model card describes it as built on Gemini 3.6 Flash with algorithmic improvements to its core reasoning foundation, aimed at complex coding, agentic workflows, and reliable multi-step execution.
How much does Gemini 3.7 Flash cost?
Google lists $0.75 per million input tokens and $3.75 per million output tokens (thinking tokens included) through 31 December 2026, rising to $1.50 and $7.50 from 1 January 2027. Context caching is $0.075 per million tokens through 2026 ($0.15 after), with cache storage at $0.50 per million tokens per hour ($1.00 after).
How does Gemini 3.7 Flash compare with Gemini 3.6 Flash?
In Google's own launch table it improves on the previous Flash model across coding and agentic work: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 48.6%, Terminal-bench 2.1 85.8% vs 78.0%, AutomationBench 30.4% vs 17.0%, and GDM-MRCR v2 (8-needle) 97.0% vs 91.8%. Google reports the same introductory price for both.
What context window and modalities does it support?
Gemini 3.7 Flash accepts text, images, audio and video with a 1 million-token context window and returns up to 64K output tokens. Its knowledge cutoff is March 2026, with some domains limited to January 2025.
Where can I use Gemini 3.7 Flash?
Google lists the Gemini API in Google AI Studio, Google Antigravity, Android Studio, the Gemini Enterprise app and Agent Platform, and the Gemini app's Spark experience for AI Pro and Ultra subscribers.