Hark launches Handoff agent, claims top benchmark score but skips latest models

Hark, a secretive AI startup founded by serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously navigates websites to complete tasks like booking flights or ordering food. The company claims top performance on the Online-Mind2Web benchmark with a 97.7 score, significantly outpacing GPT 5.4, Claude Opus 4.8, and Gemini 2.5 Pro, while pricing at less than one-tenth the token cost of competing models. Public sign-ups opened today with availability planned for later this month, though critical questions remain about how Handoff performs against the current generation of frontier models.
TL;DR
- Hark announced Handoff, a computer use agent that autonomously completes web-based tasks by controlling a dedicated virtual browser, file system, and terminal
- Handoff scored 97.7 on the Online-Mind2Web benchmark versus 92.8 for GPT 5.4, 84.1 for Claude Opus 4.8, and 69 for Gemini 2.5 Pro
- Pricing is $0.18 per million input tokens and $2.37 per million output tokens, less than one-tenth the cost of GPT 5.5, with 0.8-second per-turn latency
- Benchmark comparisons exclude current-generation models like GPT-5.6, Opus 5, DeepSeek V4, and Kimi K3, which have not published Online-Mind2Web results
Why It Matters
Computer use agents represent a significant shift in how AI systems interact with digital infrastructure. Most websites lack public APIs, forcing agents to navigate user interfaces directly, a capability that could automate substantial portions of knowledge work. Hark's claims of superior performance and lower cost suggest meaningful progress in this category, though the absence of comparisons against current frontier models limits the ability to assess true competitive standing.
Business Impact
For enterprises, computer use agents could reduce operational costs by automating routine web-based tasks across recruiting, customer service, and administrative functions. Hark's pricing model and claimed performance advantages position it as a potential alternative to building custom automation on expensive frontier models, though validation against the latest competing systems is needed before large-scale adoption decisions.
Key Implications
- The lack of benchmark comparisons against GPT-5.6, Opus 5, and strong open-source models like DeepSeek V4 makes it impossible to independently verify Hark's 'top-ever' claim
- Latency measurements were conducted by Hark using its own harness with competing models set to their slowest reasoning levels, limiting the reliability of performance comparisons
- On WebTailBench v2, one of Hark's own chosen benchmarks, GPT 5.5 outperformed Handoff (72.3 versus 68.6), suggesting the 'best' framing requires qualification
- The research finding that fewer than 1 in 1000 websites have publicly accessible APIs validates the market need for visual web navigation agents
What to Watch
Monitor whether Hark publishes benchmark results against current-generation models like GPT-5.6 and Opus 5, which have shown substantial gains in computer use tasks. Track real-world deployment outcomes and enterprise adoption rates, particularly for recruiting and customer service use cases. Watch for independent third-party evaluations of Handoff's performance and latency claims using standardized testing conditions.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.


