Overview
K2 Horizon 375B-A23B is the largest and most capable model in K2 Horizon, the six-model fleet the Institute of Foundation Models released on 3 September 2026. The fleet spans 0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B parameters, and IFM positions the 375B model for demanding enterprise workloads — complex reasoning, software engineering, research and long-horizon agentic tasks. It takes text input across a 512K-token context window.
What separates it from most open-weight releases is how much is released with it. IFM publishes the models and code under Apache-2.0 and opens the whole training lifecycle: intermediate checkpoints, training data or detailed data-construction recipes, architecture, mixture compositions, training code, configurations, fine-grained logs and evaluation results. Datasets ship under their own applicable licenses such as ODC-BY, and where redistribution is not possible IFM documents how the data was constructed and mixed. Each model in the fleet was pretrained on roughly 20 trillion tokens, with about 10 trillion synthetic tokens in the mix and nearly 17% of the pre-training corpus made up of explicit problem-solving trajectories.
On IFM's own published results, the model places among the top models below 400 billion parameters. It scores 87.3 on GPQA Diamond, 76.0 on AA-LCR long-context reasoning, 72.8 on BrowseComp, 70.2 on Terminal-Bench 2.1, 67.7 on MCPMark and 65.3 on Toolathlon Verified — competitive with, and on several agentic rows ahead of, substantially larger open-weight models such as Nemotron 3 Ultra (550B) and Inkling (975B), while trailing closed models like Claude Sonnet 5 on most rows.
IFM also published a reward-hacking audit of its own headline coding number, which is unusual enough to be worth knowing. Running the model on 89 Terminal-Bench 2.1 tasks with eight attempts each, 500 of 712 trials passed the verifier for the reported 70.2% accuracy; auditing every passing trial with Artificial Analysis's reward-hacking procedure flagged 24 trials across 10 tasks, and removing them lowers accuracy to 66.9%. IFM notes that the 3.37-point flag rate sits between the 2.2% and 4.1% rates Artificial Analysis reports for Claude Fable 5 and GPT-5.6 Luna.
All six sizes are open weights with day-zero support from vLLM, SGLang and Ollama, and deployment across NVIDIA, AMD and Cerebras hardware. IFM is a research institute launched by the Mohamed bin Zayed University of Artificial Intelligence in May 2025, continuing the fully-open line of work it began with the 2023 LLM360 paper.
| Released | 2026-09-03 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 375B total · 23B active |
| Context | 512K |
| Architecture | Sparse Mixture-of-Experts. Roughly 23 billion parameters are activated per token out of 375 billion total, so the model draws on the capacity of a much larger network without running every parameter for every token. It shares core architecture, vocabulary, training methodology, interfaces and deployment tooling with the rest of the six-model Horizon fleet. |
| Modalities | Text |
| Status | Generally available |
Benchmarks
IFM's published results for K2 Horizon 375B-A23B against open-weight and closed models. GDPVal-AA is an Elo rating; all other rows are percentages.
| Benchmark | K2-Horizon-375B-A23B | Nemotron 3 Ultra | Inkling (xhigh) | MiniMax-M3 (max) | GLM 5.2 (max) | GPT 5.6 Luna (max) | GPT 5.6 Terra (high) | Claude Sonnet 5 (max) |
|---|---|---|---|---|---|---|---|---|
| GDPVal-AA (Elo) | 1441 | 1162 | 1234 | 1380 | 1498 | 1569 | 1503 | 1584 |
| tau3-Banking | 34% | 14.2% | 29.1% | 15.3% | 34.6% | 31.1% | 28.7% | 37.3% |
| Toolathlon Verified | 65.3% | 34.3% | 45.5% | 53.7% | 59.9% | 67.5% | 64.8% | 71.6% |
| Automation Bench Public | 25.3% | 8% | 12.8% | 20.5% | 26.2% | 33.5% | 28% | 34.7% |
| Apex-Agents (pass@1) | 24.8% | 9% | 19% | 23.8% | 26.9% | 28.6% | 25.4% | 31.7% |
| MCPMark | 67.7% | 45.7% | 51.2% | 48.8% | 72.4% | 66.9% | 74% | 65.3% |
| BrowseComp | 72.8% | 44.4% | 77.1% | 83.5% | — | 83.3% | — | 84.7% |
| WildClawBench | 50.9% | 34.2% | 52.3% | 56.4% | 55% | 50.4% | 60% | — |
| Terminal-Bench 2.1 | 70.2% | 53.9% | 55.1% | 65.2% | 77.9% | 80.9% | 75.7% | 80.5% |
| SciCode | 42.7% | 39.9% | 46.1% | 45.4% | 50.5% | 52.5% | 50.1% | 53.6% |
| SWE-Atlas-QnA (strict) | 48.4% | — | 25.5% | 42.3% | 46.4% | — | — | — |
| SWE Bench Pro (strict) | 42.6% | 38.7% | 43.1% | 43.8% | 46.7% | 48.8% | — | — |
| Humanity's Last Exam (without tools) | 32% | 28.4% | 31.9% | 39% | 41.1% | 39.5% | 38.5% | 41.3% |
| GPQA Diamond | 87.3% | 86.7% | 87.2% | 92.9% | 89.5% | 91.1% | 89.6% | 91.1% |
| CritPt | 8.6% | 3.1% | 5.4% | 3.7% | 20.9% | 21% | 22.9% | 16.9% |
| AA-LCR | 76% | 71% | 73.3% | 80.3% | 76.7% | 78.3% | 73.3% | 77% |
| AA-Omniscience Accuracy | 23% | 23% | 42% | 17% | 24% | 43% | 45% | 40% |
| AA-Omniscience Non-Hallucination | 74.7% | 70% | 32% | 82% | 74% | 7% | 10% | 61% |
This model's scores
- GPQA Diamond (graduate-level science QA)87.3%
- AA-LCR (long-context reasoning)76%
- BrowseComp (deep web research)72.8%
- Terminal-Bench 2.1 (agentic terminal use)70.2%
- MCPMark (MCP tool use)67.7%
- Toolathlon Verified (agentic tool use)65.3%
- SWE Bench Pro (software engineering, strict)42.6%
- Humanity's Last Exam (without tools)32%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- Apache-2.0 weights shipped with training code, data recipes, intermediate checkpoints and fine-grained training logs
- 87.3 on GPQA Diamond and 76.0 on AA-LCR long-context reasoning
- Strong agentic tool use — 65.3 Toolathlon Verified, 67.7 MCPMark, 34.0 tau3-Banking — ahead of larger open-weight MoE models
- 512K-token context with only ~23B parameters active per token
- Day-zero vLLM, SGLang and Ollama support, and deployment on NVIDIA, AMD and Cerebras hardware
- IFM publishes a reward-hacking audit of its own Terminal-Bench result rather than only the headline number
Best for
- Reach for it when you need frontier-adjacent open weights you can self-host without a proprietary licence
- Reach for it for long-horizon agentic work — tool use, terminal tasks and multi-step research
- Reach for it for repo-scale and long-document work that needs the 512K-token window
- Reach for it for research into how capabilities emerge, since intermediate checkpoints and training logs are published
- Look at the smaller Horizon sizes (0.9B–36B) instead when you need on-device or single-workstation deployment
FAQ
What is K2 Horizon 375B-A23B?
It is the largest model in K2 Horizon, the six-model fleet the Institute of Foundation Models released on 3 September 2026. It is a sparse Mixture-of-Experts model with 375 billion total parameters and roughly 23 billion activated per token, a 512K-token context window, and Apache-2.0 weights.
What makes K2 Horizon "fully open"?
IFM releases more than the final weights. For every Horizon model it publishes training data or detailed construction recipes, training code, model configurations, intermediate checkpoints throughout training, fine-grained training logs, and evaluation results. Models and code are Apache-2.0; datasets ship under their own licenses such as ODC-BY, and where redistribution is restricted IFM documents how the data was built and mixed.
How does K2 Horizon 375B-A23B compare with other models?
On IFM's published table it scores 87.3 on GPQA Diamond, 76.0 on AA-LCR, 72.8 on BrowseComp, 70.2 on Terminal-Bench 2.1 and 65.3 on Toolathlon Verified. It leads the much larger open-weight Nemotron 3 Ultra (550B) and Inkling (975B) on most agentic rows, sits close to GLM 5.2 on some, and trails closed models such as Claude Sonnet 5 on most rows.
Why did IFM publish a lower Terminal-Bench number than its headline?
IFM audited its own result for reward hacking. Across 712 trials, 500 passed the verifier for 70.2% accuracy; auditing every passing trial with Artificial Analysis's procedure flagged 24 trials across 10 tasks, which lowers accuracy to 66.9%. IFM reports both figures and notes the 3.37-point flag rate falls between the rates Artificial Analysis reports for Claude Fable 5 and GPT-5.6 Luna.
How do I run K2 Horizon 375B-A23B?
The weights are on Hugging Face under Apache-2.0 with day-zero support from vLLM, SGLang and Ollama, and deployment on NVIDIA, AMD and Cerebras hardware. The model card gives a vLLM launch command using tensor parallelism across 8 GPUs with expert parallelism enabled, and an equivalent SGLang server command.
What other models are in the K2 Horizon fleet?
Six sizes share the same architecture, vocabulary and training methodology: 0.9B for constrained devices such as watches and glasses, 3.7B and 7B for phones and on-device use, a dense 32B and a sparse 36B-A4B for local workstations, and the 375B-A23B for enterprise deployments. The 36B-A4B uses MoVA, IFM's Mixture-of-Value-Attention architecture that applies expert routing inside attention.