September 28, 2026
Run a GUI AI agent locally: Holo4-35B-A3B with a 262K context window
H Company released Holo4-35B-A3B on Sep 28: of its 35 billion parameters, 3 billion are activated for each request. The dense Holo4-27B launched alongside it, and both models are available through the H Models API. The 35B-A3B API costs $0.30 per million input tokens and $2 per million output tokens; the 27B costs $0.40 and $3.

Holo4 can work with GUIs, code, MCP, and APIs across the web, computers, and Android. For the agent loop, H Company offers a harness: it passes the model screenshots and tool results, then executes clicks, text input, and tool calls.
Local deployment. Holo4-35B-A3B-FP8 runs through vLLM 0.28+ with `vllm serve Hcompany/Holo4-35B-A3B-FP8`; the server listens on port 8000 and accepts contexts of up to 262,144 tokens. BF16, FP8, and NVFP4 are designed for vLLM, while Q4 GGUF is for llama.cpp. The 35B-A3B weights are distributed under Apache 2.0, while the 27B, marked Research only, is under CC BY-NC 4.0.
Via API. You need an HAI_API_KEY and an OpenAI-compatible client pointed at `https://api.hcompany.ai/v1`. Holo4 requires paid credits; the free tier currently provides holo3-1-35b-a3b.
In H Company benchmarks, the 27B scored 85.2% on OSWorld at $0.08 per task, while the 35B-A3B scored 80.8% at $0.05. On 120 held-out AutomationBench tasks, results were lower: 49.3% for the 27B and 31.7% for the 35B-A3B.
On AutomationBench, 480 of 600 tasks overlap with data H Company used to assemble its training split.
