A computer-use agent you can download

H Company, the Paris-based AI lab formerly known as Hugging Face's early deployment partner and now better known simply as H, released Holo4 on September 28, 2026: a pair of computer-use agent models that click, type, run code and call business APIs, published as downloadable open weights on Hugging Face. The release matters less for any single benchmark number than for a structural shift. A category of capability that until recently lived only behind paid API keys from closed labs is now inspectable: anyone can download the weights, run the model, and replay step-by-step records of how the company's own benchmark results were produced.

The company published two sizes: Holo4-27B, a 27 billion parameter dense model, and Holo4-35B-A3B, a 35 billion parameter Mixture of Experts model with roughly 3 billion parameters active per step. It also shipped Holotron4 Nano, a smaller agentic model built by applying the same training recipe to NVIDIA's Nemotron 3 Nano Omni base. All are available in BF16, FP8, NVFP4 and Q4 GGUF formats, meaning they can run on a range of hardware from data-center GPUs to consumer machines.

One model, every interface

What distinguishes Holo4 from many prior computer-use models is that it operates across multiple interfaces rather than one. The team writes: "Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools." That combination addresses a real gap. GUI-focused agents fail when there is no screen, and tool-calling agents fail when an application offers no API.

The blog post frames the problem plainly: "Real work is not siloed that way, and a single business task can require combining these different approaches." Holo4 runs on desktops, the web, Android, in code sandboxes and against business APIs, and the same model is used in each case, so a builder does not have to route different platforms to different models.

Both sizes improve on their Qwen bases. The 27B model is a fine-tune of Qwen3.8-27B, according to its model card, which also discloses a notable licensing detail: the Holo4-27B weights are released under CC BY-NC 4.0, a non-commercial license. That limits how businesses can use the downloadable weights and is an important caveat to the word open-weights in this case. The trajectory dataset, by contrast, is Apache 2.0.

What the scores do and do not prove

On OSWorld 2.0, an academic benchmark for controlling a desktop computer, the company reports that Holo4-27B scores 61.7 percent against 81.8 percent for Opus 5.5, and that Holo4-35B-A3B reaches 30.9 percent. The company adds: "Holo4 trails only the strongest closed models on long workflows." These are the company's own reported numbers, produced in its harness; harnesses, task releases and subsets vary across papers and leaderboards, so cross-model comparisons should be read as approximate rather than exact.

The company positions the smaller model on cost. In its cost-performance notes it states that Holo4 is priced at H Models API rates and that competitor costs are estimated from token counts of actual runs at list prices. Those cost figures are self-reported vendor estimates, not independent measurements. The blog also notes that Qwen3.6 35B-A3B costs assume cache hits at 20 percent of the input price, another assumption that a reader should treat as a modeling choice rather than a fact.

On AutomationBench, Zapier's benchmark for API-driven automation, the company says its own scores and costs were measured in its internal harness, while scores for other models come from the public set and the official leaderboard, which runs on the private set. The private-set evaluation of Holo4 is still pending. Until it appears, the AutomationBench comparison remains an apples-to-mostly-apples estimate rather than a like-for-like result.

Publishing every step: the trajectory release

Perhaps the most consequential transparency practice in the release is the trajectory dataset. H Company states: "We open-source every trajectory behind our scores on public benchmarks." The dataset card on Hugging Face counts 7,366 agent trajectories from both released sizes, each containing per-step reasoning, actions, tool results and screenshots, covering OSWorld, OSWorld 2, AndroidWorld, AutomationBench, PinchBench and Agents' Last Exam, with upstream licenses respected per benchmark.

Each index row records the task, instruction, success flag, score, duration and step count. Screenshot-level privacy handling is documented: credentials, internal hosts and personal data are masked, affected screenshots are replaced with placeholders, and a few tasks are excluded. A web viewer at trajectories.hcompany.ai lets anyone replay the runs step by step without downloading anything.

For an outside researcher this is a rare resource. Most frontier labs publish only aggregate scores; published agent trajectories make it possible to audit how a score was earned, spot harness artifacts, and study failure modes. It does not, however, prove general capability: the trajectories cover benchmark tasks, not arbitrary customer workloads.

How the training data was made

A third pillar of the release is the Agentic Task Factory, the company's internal pipeline that builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. The company says it has so far produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through both a GUI and MCP. The figure is the company's own count and the tasks are internal; the release does not publish the training task set itself.

Alongside training, the team rebuilt its harness, the loop that executes model actions and manages context over hundreds of steps, using failure feedback from OSWorld 2.0 runs. The two largest changes, according to the blog, were a reliable memory able to track hundreds of steps and a shell on the desktop machine itself. Agents were made to tag why each task failed, and engineers reviewed the fixes.

Holotron4 Nano demonstrates the recipe's portability. Applying the same post-training stack to NVIDIA's Nemotron 3 Nano Omni produced absolute percentage-point improvements on GUI workflows and in MCP, API and coding-sandbox environments. The team's claim is worth quoting: "Nothing in it is size-specific." That suggests the recipe could be reapplied as new base models arrive.

What it means for builders, and what remains uncertain

For builders and small organizations, three practical consequences stand out. First, inspection: a business considering an agentic assistant can read its trajectories and see how it behaves, rather than trusting a score in a press release. Second, cost: the company positions Holo4 far below closed frontier pricing on its API, though those figures are self-reported and depend on harness and token assumptions. Third, licensing: the non-commercial CC BY-NC 4.0 license on the 27B weights, and the per-variant license terms on the other models, mean open weights are not the same as unrestricted commercial use. Organizations should read the model card license before building a product on the download.

The honest summary is that a capable, unusually versatile computer-use model now exists in a form anyone can examine, with a benchmark trail that is unusually transparent for this field, while its headline advantage over closed frontier models rests on vendor-reported cost estimates and a pending private-set evaluation. The direction of travel, capable agents becoming inspectable and affordable for small organizations, is real. The precise size of the gap to the frontier remains to be independently confirmed.