Overview
humain-m3 is an Arabic-first large language model announced on 3 September 2026 by HUMAIN, Saudi Arabia's national AI company, at LEAP in Riyadh. It is a 428-billion-parameter mixture-of-experts model with 23 billion parameters active per token, built on the MiniMax-M3 lineage and further pre-trained on more than one trillion tokens of Arabic-native content.
The model was not trained from scratch. HUMAIN commissioned MiniMax to continue pre-training its existing M3 model on Arabic data — a Gulf state's national AI programme licensing a Chinese lab's base model rather than building one itself. HUMAIN's own comparison puts humain-m3 at an 89.37% average across seven public Arabic benchmarks, ahead of GPT-5.6 SOL at 87.30%, Claude Opus 5 at 87.34%, and the MiniMax M3 base at 80.34%, and leading on five of the seven individual tests.
Access is a research and evaluation preview on HUMAIN Node, through a playground and an OpenAI-compatible API, rather than a general launch. HUMAIN says the preview period is for checking capability, safety and alignment across Arabic dialects with early users, and that it plans to publish the weights under the MiniMax Community License once that safety work is complete.
| Released | 2026-09-03 |
|---|---|
| License | Proprietary (open weights announced under the MiniMax Community License) |
| Weights | API only |
| Parameters | 428B mixture-of-experts · 23B active per token |
| Architecture | Mixture-of-Experts |
| Modalities | Text |
| Status | Research preview |
Benchmarks
humain-m3 against the frontier models HUMAIN evaluated, as published on HUMAIN Node.
| Benchmark | humain-m3 | GPT-5.6 SOL | Claude Opus 5 | MiniMax M3 |
|---|---|---|---|---|
| AlGhafa | 86.45% | 81.54% | 83.06% | 75.31% |
| ArabicMMLU | 90.7% | 88.82% | 88.77% | 81.8% |
| Arabic EXAMS | 67.67% | 66.4% | 64.8% | 61.27% |
| MadinahQA | 95.44% | 94.48% | 94.78% | 87.1% |
| AraTrust | 97.53% | 93.42% | 91.38% | 90.6% |
| ALRAGE | 94.63% | 94.91% | 94.85% | 79.32% |
| Translated MMLU | 93.2% | 91.48% | 93.78% | 87% |
| Average | 89.37% | 87.3% | 87.34% | 80.34% |
This model's scores
- ArabicMMLU90.7%
- AraTrust97.53%
- MadinahQA95.44%
- ALRAGE94.63%
- Translated MMLU93.2%
- AlGhafa86.45%
- Arabic EXAMS67.67%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- Arabic-native competence: leads five of the seven public Arabic benchmarks HUMAIN reports, with its widest margin on AraTrust (97.53% vs 93.42% for GPT-5.6 SOL)
- Large sparse capacity at modest inference cost — 428B total parameters with only 23B active per token
- More than a trillion tokens of additional Arabic-native pre-training on top of the MiniMax-M3 base, which the comparison shows lifting the base model's Arabic average from 80.34% to 89.37%
- OpenAI-compatible API on HUMAIN Node, so existing client code can point at it without a rewrite
Best for
- Building Arabic-language products where general frontier models are thinly trained on Arabic content
- Evaluating Arabic dialect coverage, safety and alignment as part of the preview programme
- Sovereign or regionally-hosted deployments that need a frontier-scale model served inside Saudi Arabia
- Comparing continued-pretraining approaches against training an Arabic model from scratch
How to access
| Provider | Model ID |
|---|---|
| HUMAIN Node ↗ | — |
FAQ
How big is humain-m3?
humain-m3 is a 428-billion-parameter mixture-of-experts model with 23 billion parameters active per token. It was further pre-trained on more than one trillion tokens of Arabic-native content on top of the MiniMax-M3 base.
Who built humain-m3?
HUMAIN, Saudi Arabia's national AI company, commissioned it from MiniMax. Rather than training from scratch, MiniMax continued pre-training its M3 model on Arabic data for HUMAIN.
Can I download the weights?
Not at announcement. HUMAIN said it plans to release the weights under the MiniMax Community License once safety training is complete; at launch the model is a research and evaluation preview only.
How do I get access?
Access is a research and evaluation preview on HUMAIN Node at node.humain.com, offered through a playground and an OpenAI-compatible API.
How does humain-m3 score on Arabic benchmarks?
On HUMAIN's own comparison across seven public Arabic benchmarks it averages 89.37%, against 87.34% for Claude Opus 5, 87.30% for GPT-5.6 SOL and 80.34% for the MiniMax M3 base, and leads on five of the seven individual tests.