AI/TLDR

Gemini 3.5 Flash-Lite

Google's July 2026 Flash-Lite — the cheapest Gemini 3 tier, faster than Gemini 3 Flash on coding and agentic tasks.

Overview

Gemini 3.5 Flash-Lite is Google's cheapest and fastest Gemini 3 model, released July 21, 2026 as the API model gemini-3.5-flash-lite. It sits under Gemini 3.6 Flash in the Gemini 3 hierarchy and is aimed at high-volume, latency-sensitive workloads — translation, classification, tagging, large-scale summarisation, and cost-sensitive agent loops.

Despite being the low end of the line, Google reports Gemini 3.5 Flash-Lite outperforms the older, larger Gemini 3 Flash on several agentic and coding evals: 54.2% vs 49.6% on SWE-Bench Pro, 74.0% vs 65.1% on OSWorld-Verified, and 54% vs 31% (Gemini 3.1 Flash-Lite) on Terminal-Bench 2.1. Artificial Analysis measured its speed at about 350 output tokens/second and its intelligence-index score at 36, both well above the median for reasoning models in its price tier.

The model accepts text, images, audio, video, and PDFs as input and returns text, with a 1 million-token context window and up to 64K output tokens. It is available through the Gemini API, Google AI Studio, Vertex AI, the Gemini app, and Search AI Overviews at $0.30 per million input tokens and $2.50 per million output tokens (with cached input at $0.03 per million). Its knowledge cutoff is March 2026.

Released2026-07-21
LicenseProprietary
WeightsAPI only
ParametersNot disclosed
Context1M
Max output64K
ArchitectureNatively multimodal Gemini 3 model in the Flash-Lite tier, optimised for high-volume, latency-sensitive workloads such as translation, classification, and large-scale summarisation, and for cheap agentic loops. Reports about 350 output tokens/second on the Artificial Analysis Index and outperforms the older Gemini 3 Flash on several coding and agentic benchmarks despite being a smaller, cheaper tier.
Knowledge cutoffMarch 2026
ModalitiesText, Vision, Audio, Video, PDF
StatusGenerally available

Benchmarks

Gemini 3.5 Flash-Lite vs the older Gemini 3 Flash and Gemini 3.1 Flash-Lite (Google launch numbers, July 21, 2026).

BenchmarkGemini 3.5 Flash-LiteGemini 3.1 Flash-LiteGemini 3 Flash
SWE-Bench Pro54.2%49.6%
OSWorld-Verified (agentic computer use)74%65.1%
Terminal-Bench 2.1 (agentic terminal coding)54%31%
GDM-MRCR v2 (long-context retrieval)72.2%60.1%
GDPval-AA v2 (economically valuable knowledge work)1140 Elo642 Elo

Comparison source ↗

This model's scores

  1. SWE-Bench Pro54.2%
  2. OSWorld-Verified (agentic computer use)74%
  3. Terminal-Bench 2.1 (agentic terminal coding)54%
  4. GDM-MRCR v2 (long-context retrieval)72.2%
  5. GDPval-AA v2 (economically valuable knowledge work)1140Elo
  6. Artificial Analysis Intelligence Index36

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.30 / 1M tokens
Cached input$0.03 / 1M tokens
Output$2.50 / 1M tokens

Standard paid tier; output includes thinking tokens. Cached input storage $1.00 per 1M tokens per hour. A free tier with usage limits is also available.

Pricing source ↗

Strengths

  • Very low price: $0.30 input / $2.50 output per million tokens, with cached input at $0.03 per million
  • About 350 output tokens/second on the Artificial Analysis Index — well above the median for reasoning models in its price tier (~107 t/s)
  • Beats the older Gemini 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%)
  • Terminal-Bench 2.1 rises to 54% (vs 31% for Gemini 3.1 Flash-Lite) — Flash-Lite is now usable for terminal-style agents
  • 1M-token multimodal context (text, image, audio, video, PDF) with 64K output

Best for

  • Reach for it as the default cheap tier for high-volume translation, classification, and tagging
  • Reach for it for cost-sensitive agent loops where you would previously have picked a Flash-tier model
  • Reach for it when you need latency below what larger Flash tiers offer
  • Reach for it for multimodal ingestion at scale — audio/video/images/PDFs cheaply within a 1M-token window

How to access

ProviderModel ID
Google Gemini API ↗gemini-3.5-flash-lite
Google Vertex AI ↗gemini-3.5-flash-lite

Gemini Flash-Lite — every version

The full lineage of the Gemini Flash-Lite line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Gemini 3.5 Flash-Litecurrent2026-07-211MProprietary
Gemini 3.1 Flash-Lite2026-03-03Proprietary
Gemini 2.5 Flash-Lite2025-06-17Proprietary
Gemini 2.0 Flash-Lite2025-02-01Proprietary
Gemini 1.5 Flash-8B2024-10-03Proprietary

FAQ

When was Gemini 3.5 Flash-Lite released?

Google released Gemini 3.5 Flash-Lite on July 21, 2026 as the API model gemini-3.5-flash-lite, alongside Gemini 3.6 Flash and the limited-access Gemini 3.5 Flash Cyber. It succeeds Gemini 3.1 Flash-Lite as the cheapest Gemini 3 tier.

What is the context window of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite has a 1 million-token input context window and can generate up to 64,000 output tokens. It accepts text, images, audio, video, and PDFs as input and returns text.

How much does Gemini 3.5 Flash-Lite cost?

On the standard paid tier of the Gemini API, Gemini 3.5 Flash-Lite is $0.30 per million input tokens and $2.50 per million output tokens, with cached input at $0.03 per million. Context-cache storage is billed at $1.00 per 1M tokens per hour. A free tier with usage limits is also available.

How does Gemini 3.5 Flash-Lite compare to Gemini 3 Flash?

Google reports it beats the older, larger Gemini 3 Flash on several evals: 54.2% vs 49.6% on SWE-Bench Pro and 74.0% vs 65.1% on OSWorld-Verified. Terminal-Bench 2.1 rose from 31% (Gemini 3.1 Flash-Lite) to 54%, and the Artificial Analysis Intelligence Index score is 36 — well above the median for reasoning models in its price tier.