A new open-weight computer-use release

H Company published Holo4 on 28 September in a team article on Hugging Face. The H Company on Hugging Face announcement presents two principal computer-use models, a 27-billion-parameter version and a 35-billion-parameter mixture-of-experts variant with roughly three billion active parameters. Their weights are available on Hugging Face in several formats, while hosted access is offered through H Models API.

This is a release by H Company, not a claim that Hugging Face developed the models. Hugging Face is the distribution platform for the weights and supporting trajectories. That matters to teams comparing hosted inference with a self-managed deployment: both routes may use the same model family, but cost, isolation and operational responsibility differ.

The models are aimed at agents that interact with graphical interfaces and software over many steps. A single screenshot answer is not the core test. The difficult work is maintaining a goal, choosing the next action, noticing an error and recovering without silently corrupting the task.

How the results are framed

H Company reports that Holo4 27B scores 61.7 per cent on OSWorld 2.0 in its evaluation, against 81.8 per cent for Claude Opus 5.5. The 35B-A3B version is reported at 30.9 per cent. Those figures illustrate both the promise of smaller open models and the remaining gap to a leading closed model in the cited setup.

The company cautions that releases, harnesses and task subsets differ across comparisons. It also says its costs are estimated from input and output tokens at the relevant API rates, while other reference points may come from public leaderboards. A chart joining those points is a guide to questions worth testing, not a controlled ranking for a particular business.

A useful transparency feature is the release of trajectories behind public benchmark scores. Developers can replay steps and inspect where an agent succeeded or failed. That is more informative than a single percentage, especially for computer use where one misplaced click can change the entire remainder of a run.

Training for long tasks

H Company describes an Agentic Task Factory that generates interactive environments and verifiable tasks from documentation, screenshots and software. It says the collection has produced around 10,000 tasks spanning web apps, desktop environments and systems with both graphical and API interfaces. This approach tries to teach agents the sequence of acting and checking, not merely answering questions about a screen.

The company also rebuilt its harness, the software loop that gives a model tools and manages context. It highlights memory across hundreds of steps and a shell available on the desktop machine. Those details are significant because benchmark performance may depend as much on the harness as on the underlying weights.

In example tasks, Holo4 works through FreeCAD modelling and a Godot game. The post includes call and token counts, showing that a successful long task can consume substantial resources. A low token rate alone may not mean a cheap completed job when an agent needs dozens of actions or retries.

Deployment choices and controls

H Company lists BF16, FP8, NVFP4 and four-bit GGUF weights on Hugging Face. These formats give implementers options for different hardware and quantisation trade-offs. An organisation should benchmark the exact format it plans to serve; a score from a higher-precision setup may not transfer cleanly to a smaller local deployment.

A computer-use agent also needs a permission design. It should work in an isolated environment for testing, with access to only the files and services required for the task. Actions that send messages, spend money or edit production records should have an explicit approval boundary. Open weights make inspection and hosting possible, but they do not supply those safeguards automatically.

For a pilot, record every action, screenshot and tool result alongside the user’s original request. Measure task completion and recovery from deliberate interruptions, not just whether the agent reached a final screen. A failure that leaves a document half-edited or a browser session authenticated can matter more than an incomplete benchmark score.

What to watch next

The post also discusses Holotron4 Nano, adapted from NVIDIA Nemotron 3 Nano Omni, and says further optimised drafter checkpoints are planned. Those are related work, but the principal Holo4 release is the subject here. Future checkpoints should be assessed on their own dates and evidence rather than assumed to share today’s performance.

The Hugging Face distribution makes the models and trajectories accessible for independent testing. It does not establish that they are suitable for any specific regulated or safety-critical workflow. Evaluators should build a representative task set, inspect failure paths and price the whole agent run, including infrastructure and review.

Holo4 adds a credible open-weight option for long computer-use experiments. Its significance is the combination of model weights, harness work and replayable evidence. The decisive question for adopters is whether that combination can complete their own work reliably under the permissions and cost limits they can support.