█

AI/TLDR

Ling-3.0-tiny

inclusionAI's lightweight hybrid reasoning model, published in August 2026: 7.9B total parameters with 1.3B active per token, MIT-licensed and built to run on a laptop.

Ling (open-weight)Open weightsGenerally available — open weights on Hugging Face
Released
2026-08
Context
262,144 tokens (with YaRN scaling, per inclusionAI's SGLang recipe)
Parameters
7.9B total · 1.3B activated per token (Mixture-of-Experts)
License
MIT

Overview

Ling-3.0-tiny is the smallest model in inclusionAI's Ling-3.0 series, from the team behind Ant Group's artificial-general-intelligence programme. It is a hybrid reasoning Mixture-of-Experts model with 7.9 billion total parameters, of which only 1.3 billion are activated per token. The weights are published on Hugging Face and ModelScope under the MIT License in BF16, FP8 and INT4 versions; the Hugging Face repositories are dated 10 August 2026, and a free endpoint was also listed on OpenRouter.

inclusionAI's Ling-3.0-tiny architecture diagram: token embeddings feed 6 groups of blocks, each stacking three KDA layers and one Gated MLA layer with RMSNorm and MoE feed-forward layers (128 experts, 8 active, plus 1 shared expert), with expanded views of the Multi-head Latent Attention and Kimi Delta Attention modules and a training objective of next-token and multi-token prediction.
The Ling-3.0-tiny architecture, from the model card.inclusionAI ↗

The model inherits the hybrid linear attention design of the Ling-3.0 series. Each 4-layer block stacks three Kimi Delta Attention (KDA) layers and one Multi-Head Latent Attention (MLA) layer, and the feed-forward layers are a sparse Mixture-of-Experts with 128 routed experts, of which 8 routed experts and 1 shared expert run for each token. Thinking is on by default and can be switched off per request through enable_thinking, so the same model serves both quick replies and multi-step reasoning.

inclusionAI built it for local and edge deployment and says it has been validated on NVIDIA DGX Spark, Apple Silicon MacBooks and the Mac mini. In FP8 it reports around 100-105 tokens per second on a DGX Spark and 86-90 tokens per second on an M4 Pro MacBook, with about 8.34 GiB peak memory at an 8K context. inclusionAI's SGLang recipe serves it on a single GPU with YaRN scaling to a 262,144-token context, and the model card also gives vLLM and Ollama (MLX on Apple Silicon) instructions.

On the model card, inclusionAI reports a score of 25 on the Artificial Analysis Intelligence Index v4.1.1 and 16 on the Artificial Analysis Agentic Index, with an output speed above 160 tokens per second in Artificial Analysis testing. Its benchmark table compares the model in thinking mode with Qwen3.5-4B, Qwen3.5-9B, Gemma-4-E4B-it and Gemma-4-12B-it, where it leads that group on GDPval v2-AA (772), TAU3-Banking-AA (20.80) and IMO-AnswerBench (71.03).

Released2026-08
LicenseMIT
WeightsOpen weights
Parameters7.9B total · 1.3B activated per token (Mixture-of-Experts)
Context262,144 tokens (with YaRN scaling, per inclusionAI's SGLang recipe)
ArchitectureHybrid-linear Mixture-of-Experts — 3 Kimi Delta Attention layers and 1 Multi-Head Latent Attention layer per 4-layer block, with 128 routed experts (8 activated) plus 1 shared expert
ModalitiesText
StatusGenerally available — open weights on Hugging Face

Benchmarks

Benchmark table comparing Ling-3.0-tiny (Thinking) with Qwen3.5-4B, Qwen3.5-9B, Gemma-4-E4B-it and Gemma-4-12B-it, all in thinking mode, across general agent, coding agent, coding, long-context, knowledge, reasoning and instruction-following benchmarks.
inclusionAI's published comparison from the Ling-3.0-tiny model card. — inclusionAI

Ling-3.0-tiny against the small models inclusionAI named in its model card, all in thinking mode. A dash means no score was given.

BenchmarkLing-3.0-tinyQwen3.5-4BQwen3.5-9BGemma-4-E4B-itGemma-4-12B-it
GDPval v2-AA772—644228645
TAU3-Banking-AA20.86.875.48.7
BFCL-v4 (FC)62.7262.4765.8341.6459.68
Terminal-Bench 2.127.725.829.21.927.3
ArtifactsBench47.9330.6138.845.1448.08
SciCode24.216.127.524.438.2
AA-LCR58.76165.33361.7
AA-Omniscience Accuracy8.5215.1216.428.5815.62
AA-Omniscience Non-Hallucination rate69.5413.4516.3969.0619.02
GPQA Diamond73.477.180.657.675.3
HLE9.39.914.93.815.7
HMMT-Feb2670.3172.8771.2133.8567.95
IMO-AnswerBench71.0363.697031.7564.09
IFBench63.615266.744.273.5
LIFEBench62.363.363.952.865.3
Multi-IF83.1579.8483.0380.9288.44

Comparison source ↗

This model's scores

  1. Artificial Analysis Intelligence Index v4.1.125
  2. IMO-AnswerBench71.03%
  3. GPQA Diamond73.4%
  4. Multi-IF83.15%
  5. BFCL-v4 (FC)62.72%
  6. Terminal-Bench 2.127.7%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Strengths

  • MIT-licensed open weights in BF16, FP8 and INT4
  • Only 1.3B of 7.9B parameters active per token, so it runs on laptops and small devices
  • 86-90 tokens/s on an M4 Pro MacBook and 100-105 tokens/s on a DGX Spark in FP8, about 8.34 GiB peak memory at 8K context (inclusionAI's figures)
  • Thinking mode that can be turned on or off per request
  • Leads Qwen3.5-9B and Gemma-4-12B-it on GDPval v2-AA and TAU3-Banking-AA in inclusionAI's table

Best for

  • Local, offline assistants and agents on a MacBook, Mac mini or DGX Spark
  • Low-cost agent steps where a small model can call tools and follow instructions
  • Edge deployment where memory is limited
  • Research on hybrid linear-attention architectures at small scale

How to access

ProviderModel ID
OpenRouter ↗inclusionai/ling-3.0-tiny:free

Ling (open-weight) — every version

The full lineage of the Ling (open-weight) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Ling-3.1-flash2026-09-30256K (trial)Not yet published
Ling-3.0-tiny2026-08—MIT
Ling-3.0-flashcurrent2026-08-02256KMIT

FAQ

Can Ling-3.0-tiny run on a laptop?

Yes. inclusionAI says it has been validated on Apple Silicon MacBooks, the Mac mini and NVIDIA DGX Spark. In FP8 it reports 86-90 tokens per second on an M4 Pro MacBook with about 8.34 GiB peak memory at an 8K context, and its Ollama instructions were verified on an M4 Pro Mac with 48 GB of memory.

How many parameters does Ling-3.0-tiny have?

7.9 billion in total, with 1.3 billion activated per token. It is a Mixture-of-Experts model with 128 routed experts, of which 8 plus 1 shared expert run for each token.

What license is Ling-3.0-tiny under?

The MIT License. BF16, FP8 and INT4 weights are on Hugging Face and ModelScope, so you can download, fine-tune and use the model commercially.

Can I turn off its thinking mode?

Yes. Thinking is on by default, and you can disable it per request with chat_template_kwargs set to enable_thinking: false. inclusionAI recommends temperature 1.0, top_p 0.95 and top_k 20.

How does Ling-3.0-tiny compare with Qwen3.5 and Gemma 4?

In inclusionAI's own table it scores higher than Qwen3.5-4B, Qwen3.5-9B, Gemma-4-E4B-it and Gemma-4-12B-it on GDPval v2-AA, TAU3-Banking-AA and IMO-AnswerBench, but lower than Qwen3.5-9B and Gemma-4-12B-it on GPQA Diamond, HLE and SciCode.