AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Inside Holo4: Building More Versatile Computer-Use Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company has released Holo4, a pair of open-weight models designed to operate software through graphical interfaces, code, MCP and APIs. The company reports a 61.7% OSWorld 2.0 score for its 27B model, but the results have not been independently verified and the two models’ scores differ sharply.

H Company has released Holo4, a series of open-weight agentic models designed to handle software tasks through graphical interfaces, code, MCP and APIs. The release includes 27B dense and 35B-A3B Mixture of Experts models, and H Company reports that the 27B version scored 61.7% on OSWorld 2.0, a benchmark for computer-use agents.

The company says Holo4 can click and type on screens, write and run code, and call MCP or API tools, selecting an interface according to the task. It is intended for use across desktops, web applications, Android, code sandboxes and business APIs. H Company presents the shared model as a way to handle work that moves between these environments without switching to a separate model for each interface.

Holo4 is available through the H Models API and as downloads on Hugging Face. H Company lists FP16, FP8 and GGUF formats. The company says the models were trained using supervised and reinforcement learning across a large collection of environments and tasks, including tasks generated with its Agentic Task Factory.

On OSWorld 2.0, H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. Its comparison gives Opus 5.5 a score of 81.8%. The company also reports results on AutomationBench for API use, measured with its internal harness, version 1.0.6. It says it has published the trajectories behind its public benchmark scores at trajectories.hcompany.ai and on Hugging Face.

At a glance
announcementWhen: Announced September 2026; independent e…
The developmentH Company released two open-weight Holo4 models for computer-use tasks spanning GUIs, code and tool calls.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

A Single Model Across Software Interfaces

Many workplace tasks combine actions that happen in different places: a person or agent may need to inspect a screen, edit a file, run code and call a business service. H Company argues that models trained for just one interface can fail when a task crosses that boundary. Holo4’s stated design aims to cover those steps with one model invoked across multiple environments.

If the reported results hold up in independent testing, open weights could give developers another option for building software automation that they can adapt or host themselves. H Company says its 27B model approaches leading closed-model performance at lower cost. That comparison remains the company’s claim; the available figures do not by themselves establish how the models perform on the same tasks, with the same evaluation setup, or in routine business use.

Publishing benchmark trajectories gives outside researchers material to inspect: they can review the sequences of actions associated with the reported results, alongside the released weights. That can make evaluation more reproducible. It does not, on its own, demonstrate that the benchmark reflects the range of conditions found in real software workflows.

From Holo Models to Holo4

Holo4 follows H Company’s earlier Holo agentic model. The announcement also describes an updated version of Holotron 3 called Holotron4 Nano. H Company’s benchmark notes identify Qwen3.8 27B as the base for the dense model and Qwen3.6 35B-A3B for the MoE model.

The company compares Holo4 with both open and closed models, but says the releases, evaluation harnesses and task subsets differ. Its cost calculations also rely on specific assumptions: Holo4 uses H Models API rates for a single run; Qwen costs use Alibaba Cloud list prices, with cache hits priced at 20% of input for the MoE model; and GPT and Opus effort sweeps draw on OpenAI launch data. Those details matter when interpreting claims that one model delivers a given score at a fraction of another model’s cost.

H Company also describes side-by-side examples using FreeCAD for 3D modeling and Godot for game design. The examples use the same prompt and harness to compare Holo4 with its Qwen base model. They are company demonstrations, while the benchmark scores offer a separate, numerical measure that still needs outside replication.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company

Benchmark Results Await Replication

The headline benchmark figures are self-reported by H Company. The company notes that evaluations of competing models may use different releases, harnesses and task subsets. For AutomationBench, it says other models’ scores come from the public set while cost figures come from a leaderboard running the private set; Holo4 has not yet been evaluated on that private set.

The reason for the large OSWorld 2.0 gap between Holo4 27B at 61.7% and Holo4 35B-A3B at 30.9% is not explained in the announcement. The reported scores also do not establish how reliably either model handles varied business workflows outside benchmark tasks and curated examples. Independent reproductions, including checks of the published trajectories, are needed to assess those questions.

Private Tests and Outside Reviews

H Company says it plans to report Holo4 results on the AutomationBench private set after that evaluation is complete. Developers and researchers can access the models through the H Models API or download them from Hugging Face, giving outside parties a way to examine the weights and attempt their own evaluations.

The next useful evidence will be comparable benchmark submissions and third-party reproductions that spell out the task sets, harnesses and costs used. Until those results are available, Holo4’s reported performance and cost advantages remain company claims rather than independently established findings.

Key Questions

What is Holo4?

Holo4 is a series of open-weight agentic models from H Company, designed to operate software through graphical interfaces, code, MCP and APIs.

Which Holo4 models are available?

H Company released a 27B dense model and a 35B-A3B Mixture of Experts model. The company lists access through its H Models API and downloads on Hugging Face.

How did Holo4 perform on OSWorld 2.0?

H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The figures have not been independently verified.

Are the benchmark results independently confirmed?

Not yet, based on the announcement. H Company says its comparisons involve different releases, harnesses and task subsets, and that Holo4 has not been evaluated on AutomationBench’s private set.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

2026’S Leading AI Motherboards For Gamers And Enthusiasts

Discover the leading AI-enabled gaming motherboards of 2026, featuring advanced connectivity, power delivery, and upgrade options for enthusiasts.

AION 2 Playtest Climbing The Steam Charts

AION 2’s latest playtest has climbed the Steam charts, reaching rank 39 with a peak of over 73,700 players, signaling growing interest.

10 Advanced AI Smartwatches To Watch Out For In 2026

Discover the 10 most advanced AI-powered smartwatches set to dominate in 2026, highlighting features, compatibility, and what makes them stand out.

2026’S Most Reliable Graphics Cards For AI And Data Science

Discover the most dependable graphics cards for AI and data science in 2026, with confirmed models, their features, and what to consider for future-proofing.