Overview
Gemini 3.5 Flash-Lite is Google's cheapest and fastest Gemini 3 model, released July 21, 2026 as the API model gemini-3.5-flash-lite. It sits under Gemini 3.6 Flash in the Gemini 3 hierarchy and is aimed at high-volume, latency-sensitive workloads — translation, classification, tagging, large-scale summarisation, and cost-sensitive agent loops.
Despite being the low end of the line, Google reports Gemini 3.5 Flash-Lite outperforms the older, larger Gemini 3 Flash on several agentic and coding evals: 54.2% vs 49.6% on SWE-Bench Pro, 74.0% vs 65.1% on OSWorld-Verified, and 54% vs 31% (Gemini 3.1 Flash-Lite) on Terminal-Bench 2.1. Artificial Analysis measured its speed at about 350 output tokens/second and its intelligence-index score at 36, both well above the median for reasoning models in its price tier.
The model accepts text, images, audio, video, and PDFs as input and returns text, with a 1 million-token context window and up to 64K output tokens. It is available through the Gemini API, Google AI Studio, Vertex AI, the Gemini app, and Search AI Overviews at $0.30 per million input tokens and $2.50 per million output tokens (with cached input at $0.03 per million). Its knowledge cutoff is March 2026.
| Released | 2026-07-21 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Parameters | Not disclosed |
| Context | 1M |
| Max output | 64K |
| Architecture | Natively multimodal Gemini 3 model in the Flash-Lite tier, optimised for high-volume, latency-sensitive workloads such as translation, classification, and large-scale summarisation, and for cheap agentic loops. Reports about 350 output tokens/second on the Artificial Analysis Index and outperforms the older Gemini 3 Flash on several coding and agentic benchmarks despite being a smaller, cheaper tier. |
| Knowledge cutoff | March 2026 |
| Modalities | Text, Vision, Audio, Video, PDF |
| Status | Generally available |
Benchmarks
Gemini 3.5 Flash-Lite vs the older Gemini 3 Flash and Gemini 3.1 Flash-Lite (Google launch numbers, July 21, 2026).
| Benchmark | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite | Gemini 3 Flash |
|---|---|---|---|
| SWE-Bench Pro | 54.2% | — | 49.6% |
| OSWorld-Verified (agentic computer use) | 74% | — | 65.1% |
| Terminal-Bench 2.1 (agentic terminal coding) | 54% | 31% | — |
| GDM-MRCR v2 (long-context retrieval) | 72.2% | 60.1% | — |
| GDPval-AA v2 (economically valuable knowledge work) | 1140 Elo | 642 Elo | — |
This model's scores
- SWE-Bench Pro54.2%
- OSWorld-Verified (agentic computer use)74%
- Terminal-Bench 2.1 (agentic terminal coding)54%
- GDM-MRCR v2 (long-context retrieval)72.2%
- GDPval-AA v2 (economically valuable knowledge work)1140Elo
- Artificial Analysis Intelligence Index36
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.30 / 1M tokens |
|---|---|
| Cached input | $0.03 / 1M tokens |
| Output | $2.50 / 1M tokens |
Standard paid tier; output includes thinking tokens. Cached input storage $1.00 per 1M tokens per hour. A free tier with usage limits is also available.
Strengths
- Very low price: $0.30 input / $2.50 output per million tokens, with cached input at $0.03 per million
- About 350 output tokens/second on the Artificial Analysis Index — well above the median for reasoning models in its price tier (~107 t/s)
- Beats the older Gemini 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%)
- Terminal-Bench 2.1 rises to 54% (vs 31% for Gemini 3.1 Flash-Lite) — Flash-Lite is now usable for terminal-style agents
- 1M-token multimodal context (text, image, audio, video, PDF) with 64K output
Best for
- Reach for it as the default cheap tier for high-volume translation, classification, and tagging
- Reach for it for cost-sensitive agent loops where you would previously have picked a Flash-tier model
- Reach for it when you need latency below what larger Flash tiers offer
- Reach for it for multimodal ingestion at scale — audio/video/images/PDFs cheaply within a 1M-token window
How to access
| Provider | Model ID |
|---|---|
| Google Gemini API ↗ | gemini-3.5-flash-lite |
| Google Vertex AI ↗ | gemini-3.5-flash-lite |
Gemini Flash-Lite — every version
The full lineage of the Gemini Flash-Lite line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Gemini 3.5 Flash-Litecurrent | 2026-07-21 | 1M | Proprietary |
| Gemini 3.1 Flash-Lite | 2026-03-03 | — | Proprietary |
| Gemini 2.5 Flash-Lite | 2025-06-17 | — | Proprietary |
| Gemini 2.0 Flash-Lite | 2025-02-01 | — | Proprietary |
| Gemini 1.5 Flash-8B | 2024-10-03 | — | Proprietary |
FAQ
When was Gemini 3.5 Flash-Lite released?
Google released Gemini 3.5 Flash-Lite on July 21, 2026 as the API model gemini-3.5-flash-lite, alongside Gemini 3.6 Flash and the limited-access Gemini 3.5 Flash Cyber. It succeeds Gemini 3.1 Flash-Lite as the cheapest Gemini 3 tier.
What is the context window of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite has a 1 million-token input context window and can generate up to 64,000 output tokens. It accepts text, images, audio, video, and PDFs as input and returns text.
How much does Gemini 3.5 Flash-Lite cost?
On the standard paid tier of the Gemini API, Gemini 3.5 Flash-Lite is $0.30 per million input tokens and $2.50 per million output tokens, with cached input at $0.03 per million. Context-cache storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.
How does Gemini 3.5 Flash-Lite compare to Gemini 3 Flash?
Google reports it beats the older, larger Gemini 3 Flash on several evals: 54.2% vs 49.6% on SWE-Bench Pro and 74.0% vs 65.1% on OSWorld-Verified. Terminal-Bench 2.1 rose from 31% (Gemini 3.1 Flash-Lite) to 54%, and the Artificial Analysis Intelligence Index score is 36 — well above the median for reasoning models in its price tier.