H Company Ships Holo4 Agents, 27B Hits 61.7% on OSWorld 2.0
TL;DR
- Holo4 ships in two sizes: a 27B dense model and a 35B-A3B mixture of experts, with weights on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF.
- On OSWorld 2.0, Holo4-27B scores 61.7% and the 35B-A3B MoE scores 30.9%; Anthropic's Opus 5.5 leads the same board at 81.8%.
- H rebuilt its execution harness with memory that spans hundreds of steps and open-sourced every trajectory behind its published benchmark scores.
H Company's new Holo4 release ships in two sizes, a 27B dense model and a 35B-A3B mixture of experts, with weights on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF. On OSWorld 2.0, the 27B scores 61.7% and the MoE scores 30.9%. Anthropic's Opus 5.5 sits at 81.8% on the same board.
"Holo4 competes with frontier models at a much lower cost per task," the H team writes. The blog post charts average partial score against cost per task but lists no dollar figure in prose.
Alongside the models, H rebuilt its execution harness, "the loop that executes the model's actions and manages its context over hundreds of steps." The biggest changes were, in H's words, "a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself."
H, a member of the NVIDIA Nemotron Coalition, also applied the same recipe to Nvidia's Nemotron 3 Nano Omni to ship Holotron4 Nano, a follow-up to Holotron 3. Every trajectory behind the published scores is open-sourced at trajectories.hcompany.ai and as a dataset on Hugging Face.
Originally reported by huggingface.co
Read the original article →Original headline: H Company Open-Weights Holo4 Computer-Use Models, 27B Scores 85.2% on OSWorld at $0.08 per Task