█

AI/TLDR

Holo4 35B-A3B

H Company's Apache-2.0 Mixture-of-Experts computer-use model, 35B total and 3B active parameters, released September 28, 2026: clicks, codes and calls tools, scoring 80.8% on OSWorld at $0.05 per task.

HoloOpen weightsGenerally available
Released
28 Sep 2026
Context
262,144 tokens (256K)
Parameters
35B total · 3B active
Input
$0.30 / 1M tokens
License
Apache-2.0

Overview

Holo4 35B-A3B is a Mixture-of-Experts vision-language model for computer use from H Company, the Paris-founded AI research lab, released on September 28, 2026 alongside the larger Holo4 27B. It has 35B total parameters with 3B active, is built on Qwen3.6-35B-A3B, and H Company describes it as close to 27B accuracy, cheaper and faster. Its weights are on Hugging Face under Apache 2.0, and it is served on H Company's OpenAI-compatible H Models API as holo4-35b-a3b, the model the API docs suggest starting with because its latency suits interactive loops.

Like Holo4 27B, it works through whatever interface a task needs: it clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, with the same call for desktop, web, Android, a code sandbox or business APIs. With H Company's hai-agents harness, screenshots and tool results go to the model and the harness executes the actions it requests.

On H Company's launch table Holo4 35B-A3B scores 80.8% on OSWorld at $0.05 per task, 77.6% on AndroidWorld at $0.07, 34.5% on the 600 public AutomationBench tasks at $0.02, and 30.9% average partial score on OSWorld 2.0 at $0.61 per task. On H Company's held-out Agentic Task Factory sets it reaches 85.4% on MCP tool servers, 70.8% on web apps and 64.8% on desktop apps. On the 120 held-out AutomationBench tasks it scores 31.7%, against 13.1% for its Qwen3.6 35B-A3B base.

H Company scatter plot of AutomationBench success rate against USD cost per task on a log scale, with Holo4 35B-A3B at about 35% near $0.02 and Holo4 27B at about 45% near $0.05, both above their Qwen base models, beside reported frontier points such as Opus 5, Fable 5, GPT-5.6 Sol and Kimi K3.
AutomationBench score against cost per task, from the Holo4 model card.H Company ↗

Training followed the Holo4 recipe: supervised fine-tuning on 127B tokens, about three quarters of them successful agentic trajectories from H Company's Agentic Task Factory (desktop 45%, web 14%, MCP and API 12%, mobile 3%), then asynchronous online reinforcement learning that trained two LoRA experts, one for desktop and web and one for terminal, MCP and API, merged back into the fine-tuned model with equal weight.

H Company diagram of the Holo4 training pipeline: 127B tokens of supervised fine-tuning (desktop 45%, web 14%, MCP and API 12%, mobile 3%, other 26%) produce a fine-tuned model; online RL trains a desktop-and-web expert and a terminal, MCP and API expert, which merge into Holo4.
Supervised fine-tuning, two RL experts, one merged model.H Company ↗

Weights come in BF16, FP8, NVFP4 and Q4 GGUF builds, with a 262,144-token maximum context in the config. On the H Models API it costs $0.30 per million input tokens ($0.03 cached) and $2.00 per million output tokens.

Released2026-09-28
LicenseApache-2.0
WeightsOpen weights
Parameters35B total · 3B active
Context262,144 tokens (256K)
ArchitectureMixture-of-Experts transformer (Qwen3.6 MoE architecture), fine-tuned from Qwen3.6-35B-A3B
ModalitiesText, Vision
StatusGenerally available

Benchmarks

H Company's Holo4 benchmark table comparing Holo4 27B, Holo4 35B-A3B and Qwen3.8 27B with frontier models (Fable 5, GPT-5.5, Qwen3.8 Max, Opus 5.5, GPT-6 Astra, Muse Spark 1.3, Opus 5, GPT-5.6 Sol, Kimi K3) on OSWorld, OSWorld 2.0, ALE-CLI, AutomationBench, AndroidWorld and Agentic Task Factory, with cost per task under each score.
H Company's published Holo4 benchmark table (September 28, 2026), including frontier reference scores and footnotes on how each was measured. — H Company

Holo4 35B-A3B against Holo4 27B and Qwen3.8 27B, transcribed from H Company's launch post (September 28, 2026). Holo4 runs are in H Company's harness; cost per task is at H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen3.8 27B. Frontier-model reference scores are in the figure above.

BenchmarkHolo4 27BHolo4 35B-A3BQwen3.8 27B
OSWorld85.2%80.8%84.3%
OSWorld — cost per task$0.08$0.05$0.22
OSWorld 2.0 (average partial score)61.7%30.9%48%
OSWorld 2.0 (success)41.5%12.3%19.4%
OSWorld 2.0 — cost per task$1.22$0.61$3.49
ALE-CLI (score)44.1%30.9%43.5%
ALE-CLI (pass rate)19.4%13.5%19%
AutomationBench (600 public tasks)45.4%34.5%40.3%
AutomationBench (120 held-out tasks)49.3%31.7%40.3%
AutomationBench — cost per task$0.05$0.02$0.09
AndroidWorld85.1%77.6%81.9%
Agentic Task Factory — Web80.2%70.8%77.6%
Agentic Task Factory — MCP89.4%85.4%74.2%
Agentic Task Factory — Desktop72%64.8%68.9%

Comparison source ↗

This model's scores

  1. Agentic Task Factory — MCP (14 tool servers)85.4%
  2. OSWorld80.8%
  3. AndroidWorld77.6%
  4. Agentic Task Factory — Web (47 web apps)70.8%
  5. Agentic Task Factory — Desktop (17 desktop apps)64.8%
  6. AutomationBench (600 public tasks)34.5%
  7. OSWorld 2.0 (average partial score)30.9%
  8. ALE-CLI (105-task Linux split)30.9%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.30 / 1M tokens
Cached input$0.03 / 1M tokens
Output$2.00 / 1M tokens

H Models API rates as listed in H Company's Holo4 launch post.

Pricing source ↗

Strengths

  • Apache 2.0 open weights, usable commercially, in BF16, FP8, NVFP4 and Q4 GGUF
  • Only 3B of 35B parameters active per token, aimed at lower cost and latency than Holo4 27B
  • 80.8% on OSWorld at $0.05 per task and 77.6% on AndroidWorld in H Company's launch table
  • 85.4% on the MCP tool-server set of H Company's held-out Agentic Task Factory
  • One model for GUI actions, code execution and MCP or API tool calls
  • 262,144-token context and an OpenAI-compatible hosted API
H Company scatter plot of OSWorld 2.0 average partial score against USD cost per task on a log scale, with Holo4 35B-A3B at about 31% near $0.61 above its Qwen3.6 35B-A3B base, and Holo4 27B at about 62% near $1.22.
OSWorld 2.0 score against cost per task, from the Holo4 model card.H Company ↗

Best for

  • Reach for it for interactive computer-use agents where latency and cost per task matter more than peak accuracy.
  • Reach for it for self-hosted, commercial agent products that need an Apache-2.0 computer-use model.
  • Reach for it for workflows that combine MCP servers and APIs with GUI steps on desktop, web or Android.
  • Look elsewhere for the longest multi-step desktop workflows: Holo4 27B scores 61.7% against 30.9% on OSWorld 2.0.

How to access

ProviderModel ID
H Models API ↗holo4-35b-a3b
Hugging Face (weights) ↗Hcompany/Holo4-35B-A3B

Holo — every version

The full lineage of the Holo line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Holo4 27Bcurrent2026-09-28256KCC-BY-NC-4.0
Holo4 35B-A3B2026-09-28256KApache-2.0

FAQ

What is Holo4 35B-A3B?

Holo4 35B-A3B is a Mixture-of-Experts vision-language model for computer use from H Company, released on September 28, 2026, with 35B total and 3B active parameters. Built on Qwen3.6-35B-A3B, it clicks and types on screens, writes and runs code, and calls MCP or API tools.

What license is Holo4 35B-A3B under?

Apache 2.0. Its base model, Qwen3.6-35B-A3B, is also Apache 2.0. The larger Holo4 27B is CC BY-NC 4.0 (non-commercial).

How well does Holo4 35B-A3B score?

In H Company's launch table it scores 80.8% on OSWorld at $0.05 per task, 77.6% on AndroidWorld, 34.5% on AutomationBench's public tasks, 30.9% on OSWorld 2.0 and 30.9% on ALE-CLI.

How much does Holo4 35B-A3B cost on the H Models API?

$0.30 per million input tokens, $0.03 per million cached input tokens and $2.00 per million output tokens. The model ID is holo4-35b-a3b.

What is the context window of Holo4 35B-A3B?

H Company lists a 256K context; the model card gives a maximum context length of 262,144 tokens in the config.

Should I use Holo4 35B-A3B or Holo4 27B?

H Company's API docs suggest starting with holo4-35b-a3b for its latency in interactive loops and moving to holo4-27b for complex multi-step tasks and novel environments. Holo4 27B scores higher on every benchmark in the launch table, most of all on OSWorld 2.0 (61.7% against 30.9%).