Overview
Holo4 35B-A3B is a Mixture-of-Experts vision-language model for computer use from H Company, the Paris-founded AI research lab, released on September 28, 2026 alongside the larger Holo4 27B. It has 35B total parameters with 3B active, is built on Qwen3.6-35B-A3B, and H Company describes it as close to 27B accuracy, cheaper and faster. Its weights are on Hugging Face under Apache 2.0, and it is served on H Company's OpenAI-compatible H Models API as holo4-35b-a3b, the model the API docs suggest starting with because its latency suits interactive loops.
Like Holo4 27B, it works through whatever interface a task needs: it clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, with the same call for desktop, web, Android, a code sandbox or business APIs. With H Company's hai-agents harness, screenshots and tool results go to the model and the harness executes the actions it requests.
On H Company's launch table Holo4 35B-A3B scores 80.8% on OSWorld at $0.05 per task, 77.6% on AndroidWorld at $0.07, 34.5% on the 600 public AutomationBench tasks at $0.02, and 30.9% average partial score on OSWorld 2.0 at $0.61 per task. On H Company's held-out Agentic Task Factory sets it reaches 85.4% on MCP tool servers, 70.8% on web apps and 64.8% on desktop apps. On the 120 held-out AutomationBench tasks it scores 31.7%, against 13.1% for its Qwen3.6 35B-A3B base.

Training followed the Holo4 recipe: supervised fine-tuning on 127B tokens, about three quarters of them successful agentic trajectories from H Company's Agentic Task Factory (desktop 45%, web 14%, MCP and API 12%, mobile 3%), then asynchronous online reinforcement learning that trained two LoRA experts, one for desktop and web and one for terminal, MCP and API, merged back into the fine-tuned model with equal weight.

Weights come in BF16, FP8, NVFP4 and Q4 GGUF builds, with a 262,144-token maximum context in the config. On the H Models API it costs $0.30 per million input tokens ($0.03 cached) and $2.00 per million output tokens.
| Released | 2026-09-28 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 35B total · 3B active |
| Context | 262,144 tokens (256K) |
| Architecture | Mixture-of-Experts transformer (Qwen3.6 MoE architecture), fine-tuned from Qwen3.6-35B-A3B |
| Modalities | Text, Vision |
| Status | Generally available |
Benchmarks

Holo4 35B-A3B against Holo4 27B and Qwen3.8 27B, transcribed from H Company's launch post (September 28, 2026). Holo4 runs are in H Company's harness; cost per task is at H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen3.8 27B. Frontier-model reference scores are in the figure above.
| Benchmark | Holo4 27B | Holo4 35B-A3B | Qwen3.8 27B |
|---|---|---|---|
| OSWorld | 85.2% | 80.8% | 84.3% |
| OSWorld — cost per task | $0.08 | $0.05 | $0.22 |
| OSWorld 2.0 (average partial score) | 61.7% | 30.9% | 48% |
| OSWorld 2.0 (success) | 41.5% | 12.3% | 19.4% |
| OSWorld 2.0 — cost per task | $1.22 | $0.61 | $3.49 |
| ALE-CLI (score) | 44.1% | 30.9% | 43.5% |
| ALE-CLI (pass rate) | 19.4% | 13.5% | 19% |
| AutomationBench (600 public tasks) | 45.4% | 34.5% | 40.3% |
| AutomationBench (120 held-out tasks) | 49.3% | 31.7% | 40.3% |
| AutomationBench — cost per task | $0.05 | $0.02 | $0.09 |
| AndroidWorld | 85.1% | 77.6% | 81.9% |
| Agentic Task Factory — Web | 80.2% | 70.8% | 77.6% |
| Agentic Task Factory — MCP | 89.4% | 85.4% | 74.2% |
| Agentic Task Factory — Desktop | 72% | 64.8% | 68.9% |
This model's scores
- Agentic Task Factory — MCP (14 tool servers)85.4%
- OSWorld80.8%
- AndroidWorld77.6%
- Agentic Task Factory — Web (47 web apps)70.8%
- Agentic Task Factory — Desktop (17 desktop apps)64.8%
- AutomationBench (600 public tasks)34.5%
- OSWorld 2.0 (average partial score)30.9%
- ALE-CLI (105-task Linux split)30.9%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.30 / 1M tokens |
|---|---|
| Cached input | $0.03 / 1M tokens |
| Output | $2.00 / 1M tokens |
H Models API rates as listed in H Company's Holo4 launch post.
Strengths
- Apache 2.0 open weights, usable commercially, in BF16, FP8, NVFP4 and Q4 GGUF
- Only 3B of 35B parameters active per token, aimed at lower cost and latency than Holo4 27B
- 80.8% on OSWorld at $0.05 per task and 77.6% on AndroidWorld in H Company's launch table
- 85.4% on the MCP tool-server set of H Company's held-out Agentic Task Factory
- One model for GUI actions, code execution and MCP or API tool calls
- 262,144-token context and an OpenAI-compatible hosted API

Best for
- Reach for it for interactive computer-use agents where latency and cost per task matter more than peak accuracy.
- Reach for it for self-hosted, commercial agent products that need an Apache-2.0 computer-use model.
- Reach for it for workflows that combine MCP servers and APIs with GUI steps on desktop, web or Android.
- Look elsewhere for the longest multi-step desktop workflows: Holo4 27B scores 61.7% against 30.9% on OSWorld 2.0.
How to access
| Provider | Model ID |
|---|---|
| H Models API ↗ | holo4-35b-a3b |
| Hugging Face (weights) ↗ | Hcompany/Holo4-35B-A3B |
Holo — every version
The full lineage of the Holo line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Holo4 27Bcurrent | 2026-09-28 | 256K | CC-BY-NC-4.0 |
| Holo4 35B-A3B | 2026-09-28 | 256K | Apache-2.0 |
FAQ
What is Holo4 35B-A3B?
Holo4 35B-A3B is a Mixture-of-Experts vision-language model for computer use from H Company, released on September 28, 2026, with 35B total and 3B active parameters. Built on Qwen3.6-35B-A3B, it clicks and types on screens, writes and runs code, and calls MCP or API tools.
What license is Holo4 35B-A3B under?
Apache 2.0. Its base model, Qwen3.6-35B-A3B, is also Apache 2.0. The larger Holo4 27B is CC BY-NC 4.0 (non-commercial).
How well does Holo4 35B-A3B score?
In H Company's launch table it scores 80.8% on OSWorld at $0.05 per task, 77.6% on AndroidWorld, 34.5% on AutomationBench's public tasks, 30.9% on OSWorld 2.0 and 30.9% on ALE-CLI.
How much does Holo4 35B-A3B cost on the H Models API?
$0.30 per million input tokens, $0.03 per million cached input tokens and $2.00 per million output tokens. The model ID is holo4-35b-a3b.
What is the context window of Holo4 35B-A3B?
H Company lists a 256K context; the model card gives a maximum context length of 262,144 tokens in the config.
Should I use Holo4 35B-A3B or Holo4 27B?
H Company's API docs suggest starting with holo4-35b-a3b for its latency in interactive loops and moving to holo4-27b for complex multi-step tasks and novel environments. Holo4 27B scores higher on every benchmark in the launch table, most of all on OSWorld 2.0 (61.7% against 30.9%).
