█

AI/TLDR

Holo4 27B

H Company's 27B dense computer-use model, released September 28, 2026: one vision-language model that clicks, writes code and calls MCP or API tools, scoring 85.2% on OSWorld at $0.08 per task.

HoloOpen weightsGenerally available
Released
28 Sep 2026
Context
262,144 tokens (256K)
Parameters
27B (dense)
Input
$0.40 / 1M tokens
License
CC-BY-NC-4.0

Overview

Holo4 27B is a vision-language model for computer use from H Company, the Paris-founded AI research lab, released on September 28, 2026 together with the smaller Holo4 35B-A3B. It is the larger of the two Holo4 sizes, a 27B dense model built on Qwen3.8-27B, and H Company positions it for the best accuracy on long, multi-step tasks across web, desktop and mobile. Its weights are on Hugging Face under the non-commercial CC BY-NC 4.0 license, and it is served on H Company's OpenAI-compatible H Models API as holo4-27b.

The model works through whatever interface a task needs: it clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, all through the same call. H Company says it runs on desktops, on the web, on Android, in a code sandbox and against business APIs, without picking a different model per platform. With H Company's hai-agents harness, screenshots and tool results go to the model and the harness executes the clicks, typing, code and tool calls it requests.

FreeCAD window showing a 3D Eiffel Tower model with four curved legs, three square platforms and a mast, and a model tree listing Leg1 to Leg4, Platform1 to Platform3 and Mast.
Holo4 27B building a replica of the Eiffel Tower in FreeCAD from a text prompt.H Company ↗

On H Company's launch table Holo4 27B scores 85.2% on OSWorld at $0.08 per task, 61.7% average partial score on the long workflows of OSWorld 2.0 (41.5% success, $1.22 per task), 85.1% on AndroidWorld, 45.4% on the 600 public AutomationBench tasks and 44.1% on ALE-CLI. In each case it improves on its Qwen3.8 27B base model. H Company says Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Opus 5.5 scores 81.8% at $8.48 per task. H Company publishes every trajectory behind its public benchmark scores.

Training started with supervised fine-tuning on 127B tokens, about three quarters of them successful agentic trajectories from H Company's Agentic Task Factory: desktop (45%), web (14%), MCP and API (12%) and mobile (3%), with the rest covering multimodal reasoning, GUI grounding and text-only tool use and coding. Asynchronous online reinforcement learning on long-horizon tasks then trained two LoRA experts, one for desktop and web and one for terminal, MCP and API, which were merged back into the fine-tuned model with equal weight and no further training.

H Company diagram of the Holo4 training pipeline: 127B tokens of supervised fine-tuning (desktop 45%, web 14%, MCP and API 12%, mobile 3%, other 26%) produce a fine-tuned model; online RL trains a desktop-and-web expert and a terminal, MCP and API expert, which merge into Holo4.
Supervised fine-tuning, two RL experts, one merged model.H Company ↗

Weights come in BF16, FP8, NVFP4 and Q4 GGUF builds, and the config sets a maximum context of 262,144 tokens. On the H Models API, Holo4 27B costs $0.40 per million input tokens ($0.04 cached) and $3.00 per million output tokens.

Released2026-09-28
LicenseCC-BY-NC-4.0
WeightsOpen weights
Parameters27B (dense)
Context262,144 tokens (256K)
ArchitectureDense transformer (Qwen3.8 dense architecture), fine-tuned from Qwen3.8-27B
ModalitiesText, Vision
StatusGenerally available

Benchmarks

H Company's Holo4 benchmark table comparing Holo4 27B, Holo4 35B-A3B and Qwen3.8 27B with frontier models (Fable 5, GPT-5.5, Qwen3.8 Max, Opus 5.5, GPT-6 Astra, Muse Spark 1.3, Opus 5, GPT-5.6 Sol, Kimi K3) on OSWorld, OSWorld 2.0, ALE-CLI, AutomationBench, AndroidWorld and Agentic Task Factory, with cost per task under each score.
H Company's published Holo4 benchmark table (September 28, 2026), including frontier reference scores and footnotes on how each was measured. — H Company

Holo4 27B against Holo4 35B-A3B and its Qwen3.8 27B base, transcribed from H Company's launch post (September 28, 2026). Holo4 runs are in H Company's harness; cost per task is at H Models API rates for Holo4 and Alibaba Cloud list prices for Qwen3.8 27B. Frontier-model reference scores are in the figure above.

BenchmarkHolo4 27BHolo4 35B-A3BQwen3.8 27B
OSWorld85.2%80.8%84.3%
OSWorld — cost per task$0.08$0.05$0.22
OSWorld 2.0 (average partial score)61.7%30.9%48%
OSWorld 2.0 (success)41.5%12.3%19.4%
OSWorld 2.0 — cost per task$1.22$0.61$3.49
ALE-CLI (score)44.1%30.9%43.5%
ALE-CLI (pass rate)19.4%13.5%19%
AutomationBench (600 public tasks)45.4%34.5%40.3%
AutomationBench (120 held-out tasks)49.3%31.7%40.3%
AutomationBench — cost per task$0.05$0.02$0.09
AndroidWorld85.1%77.6%81.9%
Agentic Task Factory — Web80.2%70.8%77.6%
Agentic Task Factory — MCP89.4%85.4%74.2%
Agentic Task Factory — Desktop72%64.8%68.9%

Comparison source ↗

This model's scores

  1. Agentic Task Factory — MCP (14 tool servers)89.4%
  2. OSWorld85.2%
  3. AndroidWorld85.1%
  4. Agentic Task Factory — Web (47 web apps)80.2%
  5. Agentic Task Factory — Desktop (17 desktop apps)72%
  6. OSWorld 2.0 (average partial score)61.7%
  7. AutomationBench (600 public tasks)45.4%
  8. ALE-CLI (105-task Linux split)44.1%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.40 / 1M tokens
Cached input$0.04 / 1M tokens
Output$3.00 / 1M tokens

H Models API rates as listed in H Company's Holo4 launch post.

Pricing source ↗

Strengths

  • One model for GUI actions, code execution and MCP or API tool calls, on desktop, web and Android
  • 85.2% on OSWorld at $0.08 per task and 85.1% on AndroidWorld in H Company's launch table
  • 61.7% average partial score on OSWorld 2.0 long workflows at $1.22 per task
  • Improves on its Qwen3.8 27B base across every benchmark H Company reported
  • Open weights in BF16, FP8, NVFP4 and Q4 GGUF, plus a hosted OpenAI-compatible API
  • 262,144-token context for agent runs of hundreds of steps
H Company scatter plot of OSWorld 2.0 average partial score against USD cost per task on a log scale, with Holo4 27B at about 62% near $1.22 and Holo4 35B-A3B at about 31% near $0.61, each above its Qwen base model, against a closed-frontier line led by GPT-6 Astra.
OSWorld 2.0 score against cost per task, from the Holo4 model card.H Company ↗

Best for

  • Reach for it for computer-use agents that operate desktop software, websites and Android apps through screenshots and clicks.
  • Reach for it for business workflows that mix GUI steps with MCP servers, APIs and code in the same run.
  • Reach for it for research and evaluation of open computer-use agents, using the published trajectories.
  • Look elsewhere for commercial self-hosting: the 27B weights are CC BY-NC 4.0 (non-commercial); Holo4 35B-A3B is Apache 2.0, or use the H Models API.

How to access

ProviderModel ID
H Models API ↗holo4-27b
Hugging Face (weights) ↗Hcompany/Holo4-27B

Holo — every version

The full lineage of the Holo line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Holo4 27Bcurrent2026-09-28256KCC-BY-NC-4.0
Holo4 35B-A3B2026-09-28256KApache-2.0

FAQ

What is Holo4 27B?

Holo4 27B is a 27B dense vision-language model for computer use from H Company, released on September 28, 2026. Built on Qwen3.8-27B, it clicks and types on screens, writes and runs code, and calls MCP or API tools across desktop, web and Android.

Is Holo4 27B open source?

Its weights are open on Hugging Face (BF16, FP8, NVFP4 and Q4 GGUF) under the CC BY-NC 4.0 license, which is non-commercial. The smaller Holo4 35B-A3B is released under Apache 2.0.

How well does Holo4 27B score on OSWorld?

H Company reports 85.2% on OSWorld at $0.08 per task, against 84.3% for its Qwen3.8 27B base, and 61.7% average partial score on OSWorld 2.0 at $1.22 per task, against 48.0% for Qwen3.8 27B and 81.8% for Opus 5.5.

How much does Holo4 27B cost on the H Models API?

$0.40 per million input tokens, $0.04 per million cached input tokens and $3.00 per million output tokens. The model ID is holo4-27b.

What is the context window of Holo4 27B?

H Company lists a 256K context; the model card gives a maximum context length of 262,144 tokens in the config.

How does Holo4 27B differ from Holo4 35B-A3B?

Holo4 27B is dense and scores higher (85.2% against 80.8% on OSWorld, 61.7% against 30.9% on OSWorld 2.0). Holo4 35B-A3B is a Mixture-of-Experts model with 3B active parameters that H Company describes as close to 27B accuracy, cheaper and faster, and it is Apache 2.0 licensed.