Overview
Ling-3.0-tiny is the smallest model in inclusionAI's Ling-3.0 series, from the team behind Ant Group's artificial-general-intelligence programme. It is a hybrid reasoning Mixture-of-Experts model with 7.9 billion total parameters, of which only 1.3 billion are activated per token. The weights are published on Hugging Face and ModelScope under the MIT License in BF16, FP8 and INT4 versions; the Hugging Face repositories are dated 10 August 2026, and a free endpoint was also listed on OpenRouter.

The model inherits the hybrid linear attention design of the Ling-3.0 series. Each 4-layer block stacks three Kimi Delta Attention (KDA) layers and one Multi-Head Latent Attention (MLA) layer, and the feed-forward layers are a sparse Mixture-of-Experts with 128 routed experts, of which 8 routed experts and 1 shared expert run for each token. Thinking is on by default and can be switched off per request through enable_thinking, so the same model serves both quick replies and multi-step reasoning.
inclusionAI built it for local and edge deployment and says it has been validated on NVIDIA DGX Spark, Apple Silicon MacBooks and the Mac mini. In FP8 it reports around 100-105 tokens per second on a DGX Spark and 86-90 tokens per second on an M4 Pro MacBook, with about 8.34 GiB peak memory at an 8K context. inclusionAI's SGLang recipe serves it on a single GPU with YaRN scaling to a 262,144-token context, and the model card also gives vLLM and Ollama (MLX on Apple Silicon) instructions.
On the model card, inclusionAI reports a score of 25 on the Artificial Analysis Intelligence Index v4.1.1 and 16 on the Artificial Analysis Agentic Index, with an output speed above 160 tokens per second in Artificial Analysis testing. Its benchmark table compares the model in thinking mode with Qwen3.5-4B, Qwen3.5-9B, Gemma-4-E4B-it and Gemma-4-12B-it, where it leads that group on GDPval v2-AA (772), TAU3-Banking-AA (20.80) and IMO-AnswerBench (71.03).
| Released | 2026-08 |
|---|---|
| License | MIT |
| Weights | Open weights |
| Parameters | 7.9B total · 1.3B activated per token (Mixture-of-Experts) |
| Context | 262,144 tokens (with YaRN scaling, per inclusionAI's SGLang recipe) |
| Architecture | Hybrid-linear Mixture-of-Experts — 3 Kimi Delta Attention layers and 1 Multi-Head Latent Attention layer per 4-layer block, with 128 routed experts (8 activated) plus 1 shared expert |
| Modalities | Text |
| Status | Generally available — open weights on Hugging Face |
Benchmarks

Ling-3.0-tiny against the small models inclusionAI named in its model card, all in thinking mode. A dash means no score was given.
| Benchmark | Ling-3.0-tiny | Qwen3.5-4B | Qwen3.5-9B | Gemma-4-E4B-it | Gemma-4-12B-it |
|---|---|---|---|---|---|
| GDPval v2-AA | 772 | — | 644 | 228 | 645 |
| TAU3-Banking-AA | 20.8 | 6.8 | 7 | 5.4 | 8.7 |
| BFCL-v4 (FC) | 62.72 | 62.47 | 65.83 | 41.64 | 59.68 |
| Terminal-Bench 2.1 | 27.7 | 25.8 | 29.2 | 1.9 | 27.3 |
| ArtifactsBench | 47.93 | 30.61 | 38.8 | 45.14 | 48.08 |
| SciCode | 24.2 | 16.1 | 27.5 | 24.4 | 38.2 |
| AA-LCR | 58.7 | 61 | 65.3 | 33 | 61.7 |
| AA-Omniscience Accuracy | 8.52 | 15.12 | 16.42 | 8.58 | 15.62 |
| AA-Omniscience Non-Hallucination rate | 69.54 | 13.45 | 16.39 | 69.06 | 19.02 |
| GPQA Diamond | 73.4 | 77.1 | 80.6 | 57.6 | 75.3 |
| HLE | 9.3 | 9.9 | 14.9 | 3.8 | 15.7 |
| HMMT-Feb26 | 70.31 | 72.87 | 71.21 | 33.85 | 67.95 |
| IMO-AnswerBench | 71.03 | 63.69 | 70 | 31.75 | 64.09 |
| IFBench | 63.61 | 52 | 66.7 | 44.2 | 73.5 |
| LIFEBench | 62.3 | 63.3 | 63.9 | 52.8 | 65.3 |
| Multi-IF | 83.15 | 79.84 | 83.03 | 80.92 | 88.44 |
This model's scores
- Artificial Analysis Intelligence Index v4.1.125
- IMO-AnswerBench71.03%
- GPQA Diamond73.4%
- Multi-IF83.15%
- BFCL-v4 (FC)62.72%
- Terminal-Bench 2.127.7%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- MIT-licensed open weights in BF16, FP8 and INT4
- Only 1.3B of 7.9B parameters active per token, so it runs on laptops and small devices
- 86-90 tokens/s on an M4 Pro MacBook and 100-105 tokens/s on a DGX Spark in FP8, about 8.34 GiB peak memory at 8K context (inclusionAI's figures)
- Thinking mode that can be turned on or off per request
- Leads Qwen3.5-9B and Gemma-4-12B-it on GDPval v2-AA and TAU3-Banking-AA in inclusionAI's table
Best for
- Local, offline assistants and agents on a MacBook, Mac mini or DGX Spark
- Low-cost agent steps where a small model can call tools and follow instructions
- Edge deployment where memory is limited
- Research on hybrid linear-attention architectures at small scale
How to access
| Provider | Model ID |
|---|---|
| OpenRouter ↗ | inclusionai/ling-3.0-tiny:free |
Ling (open-weight) — every version
The full lineage of the Ling (open-weight) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Ling-3.1-flash | 2026-09-30 | 256K (trial) | Not yet published |
| Ling-3.0-tiny | 2026-08 | — | MIT |
| Ling-3.0-flashcurrent | 2026-08-02 | 256K | MIT |
FAQ
Can Ling-3.0-tiny run on a laptop?
Yes. inclusionAI says it has been validated on Apple Silicon MacBooks, the Mac mini and NVIDIA DGX Spark. In FP8 it reports 86-90 tokens per second on an M4 Pro MacBook with about 8.34 GiB peak memory at an 8K context, and its Ollama instructions were verified on an M4 Pro Mac with 48 GB of memory.
How many parameters does Ling-3.0-tiny have?
7.9 billion in total, with 1.3 billion activated per token. It is a Mixture-of-Experts model with 128 routed experts, of which 8 plus 1 shared expert run for each token.
What license is Ling-3.0-tiny under?
The MIT License. BF16, FP8 and INT4 weights are on Hugging Face and ModelScope, so you can download, fine-tune and use the model commercially.
Can I turn off its thinking mode?
Yes. Thinking is on by default, and you can disable it per request with chat_template_kwargs set to enable_thinking: false. inclusionAI recommends temperature 1.0, top_p 0.95 and top_k 20.
How does Ling-3.0-tiny compare with Qwen3.5 and Gemma 4?
In inclusionAI's own table it scores higher than Qwen3.5-4B, Qwen3.5-9B, Gemma-4-E4B-it and Gemma-4-12B-it on GDPval v2-AA, TAU3-Banking-AA and IMO-AnswerBench, but lower than Qwen3.5-9B and Gemma-4-12B-it on GPQA Diamond, HLE and SciCode.