VFF - The signal in the noise
Model ReleaseTrending

Hark launches Handoff agent, claims top benchmark score but skips latest models

Read original
Share
Hark launches Handoff agent, claims top benchmark score but skips latest models

Hark, a secretive AI startup founded by serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously navigates websites to complete tasks like booking flights or ordering food. The company claims top performance on the Online-Mind2Web benchmark with a 97.7 score, significantly outpacing GPT 5.4, Claude Opus 4.8, and Gemini 2.5 Pro, while pricing at less than one-tenth the token cost of competing models. Public sign-ups opened today with availability planned for later this month, though critical questions remain about how Handoff performs against the current generation of frontier models.

  • Hark announced Handoff, a computer use agent that autonomously completes web-based tasks by controlling a dedicated virtual browser, file system, and terminal
  • Handoff scored 97.7 on the Online-Mind2Web benchmark versus 92.8 for GPT 5.4, 84.1 for Claude Opus 4.8, and 69 for Gemini 2.5 Pro
  • Pricing is $0.18 per million input tokens and $2.37 per million output tokens, less than one-tenth the cost of GPT 5.5, with 0.8-second per-turn latency
  • Benchmark comparisons exclude current-generation models like GPT-5.6, Opus 5, DeepSeek V4, and Kimi K3, which have not published Online-Mind2Web results

Computer use agents represent a significant shift in how AI systems interact with digital infrastructure. Most websites lack public APIs, forcing agents to navigate user interfaces directly, a capability that could automate substantial portions of knowledge work. Hark's claims of superior performance and lower cost suggest meaningful progress in this category, though the absence of comparisons against current frontier models limits the ability to assess true competitive standing.

For enterprises, computer use agents could reduce operational costs by automating routine web-based tasks across recruiting, customer service, and administrative functions. Hark's pricing model and claimed performance advantages position it as a potential alternative to building custom automation on expensive frontier models, though validation against the latest competing systems is needed before large-scale adoption decisions.

  • The lack of benchmark comparisons against GPT-5.6, Opus 5, and strong open-source models like DeepSeek V4 makes it impossible to independently verify Hark's 'top-ever' claim
  • Latency measurements were conducted by Hark using its own harness with competing models set to their slowest reasoning levels, limiting the reliability of performance comparisons
  • On WebTailBench v2, one of Hark's own chosen benchmarks, GPT 5.5 outperformed Handoff (72.3 versus 68.6), suggesting the 'best' framing requires qualification
  • The research finding that fewer than 1 in 1000 websites have publicly accessible APIs validates the market need for visual web navigation agents

Monitor whether Hark publishes benchmark results against current-generation models like GPT-5.6 and Opus 5, which have shown substantial gains in computer use tasks. Track real-world deployment outcomes and enterprise adoption rates, particularly for recruiting and customer service use cases. Watch for independent third-party evaluations of Handoff's performance and latency claims using standardized testing conditions.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

AI Coding Agents Hit Cost Reality, Teams Rethink Code Review
TrendingNews

AI Coding Agents Hit Cost Reality, Teams Rethink Code Review

AI coding agents are now handling up to 99% of development work at companies like Kilo Code, forcing teams to rethink code review, cost management, and model selection. Replit, Kilo Code, and Symbotic shared strategies for deploying agentic AI safely, including risk-scoring pull requests, supporting multiple models, and capping token usage to prevent budget overruns. The shift reveals a clear divide: agents excel at greenfield development but struggle with brownfield maintenance of existing codebases, requiring human oversight at different stages.

by taryn.plumb@venturebeat.com (Taryn Plumb)· VentureBeat AI
Tech Firms Build Custom Coding Agents to Cut AI Costs
TrendingNews

Tech Firms Build Custom Coding Agents to Cut AI Costs

Coinbase, Shopify, and Ramp have built proprietary AI coding agents for internal use to reduce reliance on expensive third-party solutions from Anthropic and OpenAI. Coinbase's tool, called Forge, launched to all engineers in April 2026 and is accessible via Slack, GitHub, and a web interface. The move reflects a broader trend among well-resourced tech firms to develop custom alternatives rather than pay for commercial coding AI services.

by Laura Bratton· The Information
AI Alliance Proposes Shared Framework for Cybersecurity Incident Reporting
TrendingNews

AI Alliance Proposes Shared Framework for Cybersecurity Incident Reporting

The Open Secure AI Alliance, comprising over 120 organizations, is developing SAFE (Shared AI Findings Exchange) guidelines to standardize how agentic AI cybersecurity incidents are collected, analyzed, and shared across the ecosystem. The Linux Foundation released a Request for Comments on the framework, which proposes confidential incident collection, impact notification, control failure identification, and evidence-based recommendations to reduce systemic risk. NVIDIA, Cisco, CrowdStrike, Hugging Face, and Red Hat are among the contributors to the initial proposal, unveiled as Black Hat conference begins in Las Vegas.

by Justin Boitano· NVIDIA Blog (AI)
Alibaba's Qwen3.8-Max claims agentic AI lead, plans open-weight release

Alibaba's Qwen3.8-Max claims agentic AI lead, plans open-weight release

Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model targeting autonomous software engineering and enterprise automation. The company claims the model outperforms GPT-5.6 Sol Max and Fable 5 on agentic computing benchmarks, particularly on OSWorld-Verified (86.1 vs 83.2 and 85.0 respectively). Alibaba plans to release open weights next week, though licensing terms remain undisclosed, which could reshape enterprise adoption if permissive.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI