← All posts / Models

Click, Code, Call: H Company's Holo4 Runs the Desktop at $0.08 a Task

French lab H Company has open-weighted Holo4, a generalist computer-use agent that scores 85.2% on OSWorld at $0.08 per task and 61.7% on long-horizon OSWorld 2.0 — chasing Opus 5.5 at roughly one-seventh the cost.

Click, Code, Call: H Company's Holo4 Runs the Desktop at $0.08 a Task

On September 28, Paris-based AI lab H Company released Holo4, a new family of open-weight “computer-use” models designed to operate software the way people do — and the pricing story is as aggressive as the benchmarks. The flagship Holo4 27B scores 85.2% on OSWorld, the standard testbed for controlling a real desktop environment, at a measured cost of $0.08 per task. On the newer, much harder OSWorld 2.0 suite of long-horizon workflows, the same model reaches 61.7% — behind Anthropic’s Opus 5.5 at 81.8%, but at $1.22 per task versus $8.48. A frontier-class capability, in other words, is being sold — and in one variant given away — at a fraction of frontier pricing.

One model, every interface

The core design bet behind Holo4 is generalization across interfaces. Most agentic models today are trained for a single mode of operation: GUI-focused models are blind without a screen, while tool-calling models are stuck when an application exposes no API. H Company’s argument is that real business work is not siloed that way — a single task might require reading a PDF, clicking through a legacy ERP screen, writing a script, and calling an MCP tool.

Holo4 handles all of them with the same weights. It clicks and types on screens (desktop, web, and Android), writes and executes its own code in a sandbox, and calls MCP or REST API tools, choosing whichever interface fits the task. There is no separate model per platform and no special routing; the same checkpoint is invoked the same way whether it is driving FreeCAD on a Linux desktop or paging through a mobile app.

The numbers

H Company’s own harness puts Holo4 ahead of its open-weight base model — Alibaba’s Qwen3.8 27B — and within reach of closed frontier models on several suites:

  • OSWorld (computer tasks): Holo4 27B scores 85.2% at $0.08/task; the 35B-A3B mixture-of-experts variant scores 80.8% at $0.05. The Qwen3.8 27B base manages 84.3% but at $0.22/task. Public frontier reference points: Fable 5 at 86.0%, Qwen3.8 Max at 86.1%, GPT-5.5 at 78.7%.
  • OSWorld 2.0 (long workflows): Holo4 27B reaches 61.7% average partial score at $1.22/task. That trails Opus 5.5 (81.8%, $8.48) and GPT-6 Astra (73.5%, $9.07), but comfortably beats its own base model, Qwen3.8 27B, at 48.0% and $3.49.
  • ALE-CLI (the 105-task Linux split of Agents’ Last Exam): Holo4 27B scores 44.1% average with a 19.4% pass rate at $0.82/task, against Opus 5.5’s 63.7%/34.3% at $8.22.
  • AutomationBench (business automation via MCP): 45.4% at $0.05/task, essentially matching GPT-5.6 Sol (45.8%) and Kimi K3 (46.7%) at 8–13x lower cost.
  • AndroidWorld (phone apps): 85.1% at $0.08/task for the 27B model.

Two caveats worth stating plainly: these are vendor-run numbers (frontier comparisons are drawn from other labs’ own reports, across different harnesses and effort levels), and on pure capability the closed frontier still wins. The claim H Company is actually making is about the capability-per-dollar frontier — and there, the gap has collapsed.

How it was built

The training pipeline has three stages, and the data engine behind them is the most interesting part.

Agentic Task Factory. H Company built an internal pipeline that constructs interactive environments and verifiable tasks from documentation alone — product manuals, help centers, screenshots of real websites, and self-hosted open-source software. Coding agents then build the environment, the tasks, and their verifiers; GUI agents attempt every task, and failures feed audit-and-harden loops. A task is kept only if its verifier fails on the untouched seed state, passes on the golden state, and rejects near-miss solutions. The factory has produced roughly 10,000 tasks: about 4,000 web apps, 3,000 MCP servers, and 3,000 desktop/OS tasks, including hybrid environments that expose the same state through both a GUI and MCP.

Supervised fine-tuning. 127B tokens, roughly three-quarters of them successful agentic trajectories from the factory (desktop 45%, web 14%, MCP/API 12%, mobile 3%), with the remainder covering multimodal reasoning, GUI grounding, and text-only tool use.

Two RL experts, merged. Asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model — one for desktop and web, one for terminal, MCP, and API work. Both merge back into the base with equal weight and no further training.

The harness engineering is a story in itself. H Company rebuilt its agent loop using failure feedback from OSWorld 2.0 runs, and charts six distinct iterations — running the shell on the desktop machine itself, passing context through during memory compaction, scaling model serving, extending runs from 200 to 500 steps (2 to 6 hours) — that carried the same 27B model’s OSWorld 2.0 score from base-model territory to 61.7%. A large share of agentic “capability,” the chart quietly argues, lives in the loop around the model.

Holotron4 Nano and the Nemotron recipe test

Alongside Holo4, the company shipped Holotron4 Nano, applying the same post-training stack to Nvidia’s Nemotron 3 Nano Omni as part of the NVIDIA Nemotron Coalition. The recipe transfers strikingly well: OSWorld jumps from 21.0 to 76.3 (+55.3 points), AutomationBench from 19.4 to 35.6. Nothing in the stack is size-specific, H Company says — a claim that matters for anyone hoping to replicate it.

Availability and licensing — read the fine print

Both Holo4 sizes are served on the H Models API (27B at $0.40/M input and $3.00/M output, 35B-A3B at $0.30/$2.00, with 256K context and cached-input rates around a tenth of list). Weights are on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF, with DSpark drafter checkpoints promised in the coming days.

The licensing split deserves attention: Holo4 35B-A3B ships under Apache 2.0, but the flagship Holo4 27B weights are released under CC BY-NC 4.0 — non-commercial. Anyone planning to build a product on the best-performing checkpoint needs a commercial arrangement; the Apache-licensed MoE variant is the free path. Both derive from Alibaba’s Qwen3.8.

H Company is also publishing every trajectory behind its public benchmark scores, replayable step-by-step at trajectories.hcompany.ai and downloadable from Hugging Face — a level of scoring transparency that frontier labs have not matched.

Why it matters

Computer use is widely expected to be the first agentic domain with mass enterprise value, precisely because so much business software was never built with APIs in mind. Until now, credible performance there has been the preserve of a handful of closed frontier models. Holo4 does not close that gap — Opus 5.5 retains a 20-point lead on OSWorld 2.0 — but it demonstrates that a 27B open-weight model, trained on factory-generated tasks with a well-engineered harness, can deliver most of the practical capability at an order of magnitude lower cost, run locally, and be audited end-to-end. For the automation market, that is the difference between a demo and a deployment.