AI/TLDR

Qwen3.7-Flash

Alibaba's cheapest Qwen3.7 vision-language tier, released July 2026 — 1M-token context, text/image/video input, priced for high-volume agent traffic.

Overview

Qwen3.7-Flash is the Flash tier of Alibaba's Qwen3.7 native vision-language series, published on Alibaba Cloud Model Studio with the snapshot `qwen3.7-flash-2026-07-15` and listed on OpenRouter from July 27, 2026. Alibaba describes the 3.7 Flash models as comprehensively enhancing multimodal understanding and agent execution over 3.6-Flash, with stronger foundational multimodal ability, better object recognition, and improved real-world perception and spatial intelligence — including in multimodal agent scenarios such as Search Agent and CI Agent.

It takes text, image and video as input and returns text. The context window is 1,000,000 tokens, with a maximum input of 991,808 tokens (983,616 in thinking mode), a maximum output of 131,072 tokens, and a chain-of-thought budget of up to 262,144 tokens. Model Studio exposes function calling, structured outputs, prefix completion, context caching, batch inference and (in most regions) web search; fine-tuning is not supported.

The model is served from Model Studio in the Beijing, Singapore, Tokyo, Frankfurt and US (Virginia) regions through OpenAI-compatible, Anthropic-compatible and DashScope endpoints, and is resold through gateways such as OpenRouter. Pricing is tiered by prompt length, which makes it the cost floor of the Qwen3.7 line for high-volume multimodal and sub-agent workloads.

Released2026-07
LicenseProprietary (API-only)
WeightsAPI only
Context1M
Max output131,072 tokens
ModalitiesText, Vision, Video
StatusGenerally available

Pricing

Input¥0.2 / 1M tokens
Output¥0.8 / 1M tokens

Alibaba Cloud Model Studio list price in the Beijing region for prompts up to 32K tokens; ¥0.6 / ¥2.4 for 32K–256K and ¥1.2 / ¥4.8 for 256K–1M. Singapore-region rates are ¥0.225 / ¥0.974 at the ≤32K tier. OpenRouter lists the model at $0.03 / $0.13 per 1M tokens.

Pricing source ↗

Strengths

  • 1M-token context with text, image and video input at the cheapest tier of the Qwen3.7 line
  • Tuned for multimodal agent execution — Alibaba calls out Search Agent and CI Agent scenarios over 3.6-Flash
  • Stronger object recognition, real-world perception and spatial intelligence than the 3.6 Flash generation
  • Function calling, structured outputs, context caching and batch inference available on Model Studio
  • Served from five Model Studio regions via OpenAI-compatible, Anthropic-compatible and DashScope endpoints

Best for

  • Reach for it as the cheap sub-agent or fan-out worker behind a more expensive planner model when the work is high-volume.
  • Reach for it for screenshot, document and video understanding pipelines where per-token cost dominates the bill.
  • Reach for it for long-context multimodal retrieval and summarisation that has to stay inside one 1M-token request.
  • Reach for it for batch-mode classification and extraction over large image or video corpora.

How to access

ProviderModel ID
Alibaba Cloud Model Studio (DashScope) ↗qwen3.7-flash
OpenRouter ↗qwen/qwen3.7-flash

FAQ

What is Qwen3.7-Flash?

Qwen3.7-Flash is the Flash (cost-optimised) tier of Alibaba's Qwen3.7 native vision-language series. It accepts text, image and video input, returns text, and carries a 1,000,000-token context window. Alibaba Cloud Model Studio documents it under the snapshot qwen3.7-flash-2026-07-15, and OpenRouter listed it on July 27, 2026.

How does Qwen3.7-Flash differ from Qwen3.6-Flash?

Alibaba says the 3.7 Flash models comprehensively enhance multimodal understanding and agent execution compared with 3.6-Flash, specifically citing stronger foundational multimodal abilities, better object recognition, and improved real-world perception and spatial intelligence, with significant upgrades in multimodal agent scenarios such as Search Agent and CI Agent.

How much does Qwen3.7-Flash cost?

Model Studio prices it by prompt length. In the Beijing region it is ¥0.2 per 1M input tokens and ¥0.8 per 1M output tokens for prompts up to 32K, rising to ¥0.6 / ¥2.4 for 32K–256K and ¥1.2 / ¥4.8 for 256K–1M. The Singapore region charges ¥0.225 / ¥0.974 at the ≤32K tier. OpenRouter lists it at $0.03 / $0.13 per 1M tokens.

What are Qwen3.7-Flash's token limits?

The context window is 1,000,000 tokens. Maximum input is 991,808 tokens (983,616 in thinking mode), maximum output is 131,072 tokens, and the chain-of-thought budget goes up to 262,144 tokens.

Can Qwen3.7-Flash be fine-tuned or self-hosted?

No. It is a proprietary, API-only model — no weights are published — and Alibaba Cloud Model Studio lists fine-tuning as unsupported for it in every region. It is served from Beijing, Singapore, Tokyo, Frankfurt and US (Virginia) via OpenAI-compatible, Anthropic-compatible and DashScope endpoints.