VFF - The signal in the noise
Model ReleaseTrending

Hark launches Handoff agent, claims top benchmark score but skips latest models

Read original
Share
Hark launches Handoff agent, claims top benchmark score but skips latest models

Hark, a secretive AI startup founded by serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously navigates websites to complete tasks like booking flights or ordering food. The company claims top performance on the Online-Mind2Web benchmark with a 97.7 score, significantly outpacing GPT 5.4, Claude Opus 4.8, and Gemini 2.5 Pro, while pricing at less than one-tenth the token cost of competing models. Public sign-ups opened today with availability planned for later this month, though critical questions remain about how Handoff performs against the current generation of frontier models.

  • Hark announced Handoff, a computer use agent that autonomously completes web-based tasks by controlling a dedicated virtual browser, file system, and terminal
  • Handoff scored 97.7 on the Online-Mind2Web benchmark versus 92.8 for GPT 5.4, 84.1 for Claude Opus 4.8, and 69 for Gemini 2.5 Pro
  • Pricing is $0.18 per million input tokens and $2.37 per million output tokens, less than one-tenth the cost of GPT 5.5, with 0.8-second per-turn latency
  • Benchmark comparisons exclude current-generation models like GPT-5.6, Opus 5, DeepSeek V4, and Kimi K3, which have not published Online-Mind2Web results

Computer use agents represent a significant shift in how AI systems interact with digital infrastructure. Most websites lack public APIs, forcing agents to navigate user interfaces directly, a capability that could automate substantial portions of knowledge work. Hark's claims of superior performance and lower cost suggest meaningful progress in this category, though the absence of comparisons against current frontier models limits the ability to assess true competitive standing.

For enterprises, computer use agents could reduce operational costs by automating routine web-based tasks across recruiting, customer service, and administrative functions. Hark's pricing model and claimed performance advantages position it as a potential alternative to building custom automation on expensive frontier models, though validation against the latest competing systems is needed before large-scale adoption decisions.

  • The lack of benchmark comparisons against GPT-5.6, Opus 5, and strong open-source models like DeepSeek V4 makes it impossible to independently verify Hark's 'top-ever' claim
  • Latency measurements were conducted by Hark using its own harness with competing models set to their slowest reasoning levels, limiting the reliability of performance comparisons
  • On WebTailBench v2, one of Hark's own chosen benchmarks, GPT 5.5 outperformed Handoff (72.3 versus 68.6), suggesting the 'best' framing requires qualification
  • The research finding that fewer than 1 in 1000 websites have publicly accessible APIs validates the market need for visual web navigation agents

Monitor whether Hark publishes benchmark results against current-generation models like GPT-5.6 and Opus 5, which have shown substantial gains in computer use tasks. Track real-world deployment outcomes and enterprise adoption rates, particularly for recruiting and customer service use cases. Watch for independent third-party evaluations of Handoff's performance and latency claims using standardized testing conditions.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

AWS Launches AI Hiring Tool to Speed Recruitment at Scale
Model Release

AWS Launches AI Hiring Tool to Speed Recruitment at Scale

AWS launched Amazon Connect Talent, an AI-powered hiring platform designed to accelerate recruitment at scale for industries like retail, logistics, and hospitality. The tool uses AI agents to conduct interviews and assessments, while providing recruiters with scored candidates, transcripts, and evaluation reasoning for final decision-making. The solution aims to address bottlenecks in hiring workflows where applications pile up and strong candidates move on before contact.

by Ayesha Borker· AWS Machine Learning Blog
UN Partners With Google to Make Global Data AI-Ready
TrendingNews

UN Partners With Google to Make Global Data AI-Ready

The United Nations has partnered with Google to restructure its global development data for use by AI agents, following a UNICEF assessment that found leading AI models struggled to accurately retrieve global development statistics. The initiative addresses a critical gap where AI systems cannot reliably access or interpret UN data at scale. This partnership aims to make UN datasets machine-readable and optimized for AI-driven queries and analysis.

by Jagmeet Singh· TechCrunch AI
Instinct and Meta's Muse Add Calling to AI Agents

Instinct and Meta's Muse Add Calling to AI Agents

Two AI agent platforms, Instinct and Meta's Muse, have both added calling capabilities to their assistants. Users can now leverage these agents to perform tasks like making restaurant reservations and canceling subscriptions through voice calls. This development represents a shift toward more autonomous AI agents capable of handling real-world interactions on behalf of users.

by Ivan Mehta· TechCrunch AI
Snap launches Specs Intelligence AI assistant for iOS and Mac

Snap launches Specs Intelligence AI assistant for iOS and Mac

Snap is launching Specs Intelligence, an AI assistant designed to connect to other digital accounts and help users manage work tasks and travel information. The tool positions itself as an 'anticipatory AI service' that prioritizes daily attention items to support longer-term goals, similar to Meta's Muse and Google's Gemini Spark. It launches alongside Snap's first consumer AR glasses and is available on iOS today, with Mac support coming.

by Jay Peters· The Verge AI