Holo4 Turns One Agent Loose Across Screens, Code, and APIs
Hcompany's new Holo4 models promise a single agent that clicks, codes, and calls APIs interchangeably. Here's how an operator would actually stage-gate the switch from siloed tools to one interface-agnostic system.
Independent UpShaqo analysis built from fresh, attributed sources. We explain the impact instead of repeating the announcement.
Read for leverage: focus on the workflow change, the customer problem, and the next action—not only the product announcement.
Picture the stack most operations teams have quietly assembled over the past two years: an RPA tool that clicks through a legacy GUI, a separate coding agent that writes and runs scripts, and a third integration layer that calls whatever APIs the vendor actually exposed. Each piece works within its lane. None of them talk to each other, and every handoff between them is a person copying context from one tool into the next. Hcompany's newly released Holo4 models are built specifically to erase that seam, and the release gives operators enough detail to actually plan around it rather than just admire it.
The problem Holo4 is aimed at
Hcompany's own framing of the gap is blunt: most agentic models are trained for one interface only, so GUI-focused models are blind without a screen, while tool-calling models are stuck in front of an application that has no API. Real business tasks rarely respect that boundary. A single workflow — reconciling an invoice, provisioning an account, updating a CRM record — can require reading a screen, running a script, and hitting an internal API in the same sequence. Holo4 is designed to be the same model across all of it, called the same way whether it's running on a desktop, the web, Android, a code sandbox, or against a business API.
Meet the operator, before and after
Consider, for illustration, an operations lead responsible for a back-office process that spans a browser-based vendor portal, an internal Python script for data cleanup, and a ticketing API. Before Holo4-style models, that lead is managing three separate tools with three separate failure modes: the GUI bot breaks when the portal changes its layout, the script needs a developer to maintain it, and the API integration requires its own connector and monitoring. Each tool has its own vendor relationship, its own audit trail, and its own escalation path when something silently stalls mid-task.
After adopting an interface-agnostic model like Holo4, that same workflow is, in principle, a single agent session: it reads the portal, writes and executes the cleanup code, and files the ticket through the API, switching modes as the task demands rather than as the toolchain dictates. The operational upside isn't just fewer vendors — it's fewer places where a handoff can silently fail.
What the benchmarks actually say
Hcompany doesn't claim parity with the strongest closed models, and the numbers back that restraint. On OSWorld 2.0, a benchmark for long, multi-step desktop workflows, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, while the 35B-A3B Mixture-of-Experts variant reaches 30.9%. That's a real gap on the hardest, longest-horizon tasks. The pitch is cost, not raw ceiling: Holo4 is positioned to compete with frontier models on AutomationBench and OSWorld 2.0 at a fraction of the price per task, with Hcompany pricing Holo4 at H Models API rates and estimating cost from input and output tokens rather than claiming outright superiority.
The company also open-sourced the trajectories behind its scores, letting anyone replay each step of a benchmark run rather than take the leaderboard number on faith. For an operator, that's an unusually useful audit trail — you can inspect exactly where a model wandered off-task before you ever put it in front of a real workflow.
Efficiency isn't uniform — and that's worth knowing
Hcompany's own side-by-side demos against its Qwen base model complicate any simple "faster and cheaper" story. Building a detailed Eiffel Tower model in FreeCAD, Holo4 27B took 84 calls and 1.3M tokens versus the base model's 60 calls and 1.0M tokens — more effort, presumably in service of a more faithful result. On a Godot Pac-Man build, the pattern reverses sharply: Holo4 27B needed 68 calls and 2.4M tokens against the base model's 197 calls and 11.4M tokens, producing shorter code in the process. The lesson for operators is that efficiency gains are task-dependent, not a flat multiplier — worth measuring per workflow rather than assuming from a single demo.
An implementation sequence, not a switch flip
A disciplined rollout looks less like a swap and more like a staged evaluation:
- Access first. Both sizes are live on the H Models API, and weights are published on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF, which matters for teams that need to self-host rather than call a hosted endpoint.
- Pick a size deliberately. The 27B dense model is the cheaper entry point for testing; the 35B-A3B MoE model targets harder, longer workflows, though its OSWorld 2.0 score (30.9%) shows the ceiling is still well below frontier closed models.
- Map your own tasks to the benchmarks. Before trusting Holo4 on a live process, replay comparable trajectories from Hcompany's public set to see how the model handles multi-step recovery and error correction, the behaviors OSWorld 2.0 is designed to stress.
- Watch the harness, not just the model. Hcompany rebuilt its training harness specifically to give the agent reliable memory across hundreds of steps and shell access on the desktop machine itself — a reminder that agent quality is inseparable from the execution loop wrapped around it, something operators will need to replicate or trust in their own deployment.
- Keep an eye on Holotron4 Nano. As part of the NVIDIA Nemotron Coalition, Hcompany applied the same post-training recipe to Nemotron 3 Nano Omni, producing a lighter generalist agent. That portability signals the recipe isn't tied to one foundation model, which matters for teams wary of foundation-model lock-in.
Measuring whether it's actually working
Benchmark scores translate into a few concrete operator-facing questions: What's the task completion rate on your own long workflows, benchmarked against the 61.7%-vs-81.8% gap Holo4 shows against Opus 5.5? What's the cost per completed task once you account for retries, since AutomationBench and OSWorld 2.0 cost comparisons are estimated from token usage rather than flat pricing? And how many tool-switches — GUI to code to API — does a given workflow require, since that's precisely the seam Holo4 is built to remove.
The honest tradeoff
Hcompany is candid that releases, harnesses, and task subsets vary across the benchmarks it cites, which makes apples-to-apples comparison genuinely hard even within its own charts. Operators should treat the cost-performance advantage as directionally credible, not a precise multiplier, and should expect the capability gap on the hardest long-horizon tasks to persist until the next generation closes it further.
Sources
- Hugging Face — "Holo4: powering generalist computer-use agents"