VFF - The signal in the noise
Model ReleaseTrending

Hark launches Handoff agent, claims top benchmark score but skips latest models

Read original
Share
Hark launches Handoff agent, claims top benchmark score but skips latest models

Hark, a secretive AI startup founded by serial entrepreneur Brett Adcock, launched Handoff, a computer use agent that autonomously navigates websites to complete tasks like booking flights or ordering food. The company claims top performance on the Online-Mind2Web benchmark with a 97.7 score, significantly outpacing GPT 5.4, Claude Opus 4.8, and Gemini 2.5 Pro, while pricing at less than one-tenth the token cost of competing models. Public sign-ups opened today with availability planned for later this month, though critical questions remain about how Handoff performs against the current generation of frontier models.

  • Hark announced Handoff, a computer use agent that autonomously completes web-based tasks by controlling a dedicated virtual browser, file system, and terminal
  • Handoff scored 97.7 on the Online-Mind2Web benchmark versus 92.8 for GPT 5.4, 84.1 for Claude Opus 4.8, and 69 for Gemini 2.5 Pro
  • Pricing is $0.18 per million input tokens and $2.37 per million output tokens, less than one-tenth the cost of GPT 5.5, with 0.8-second per-turn latency
  • Benchmark comparisons exclude current-generation models like GPT-5.6, Opus 5, DeepSeek V4, and Kimi K3, which have not published Online-Mind2Web results

Computer use agents represent a significant shift in how AI systems interact with digital infrastructure. Most websites lack public APIs, forcing agents to navigate user interfaces directly, a capability that could automate substantial portions of knowledge work. Hark's claims of superior performance and lower cost suggest meaningful progress in this category, though the absence of comparisons against current frontier models limits the ability to assess true competitive standing.

For enterprises, computer use agents could reduce operational costs by automating routine web-based tasks across recruiting, customer service, and administrative functions. Hark's pricing model and claimed performance advantages position it as a potential alternative to building custom automation on expensive frontier models, though validation against the latest competing systems is needed before large-scale adoption decisions.

  • The lack of benchmark comparisons against GPT-5.6, Opus 5, and strong open-source models like DeepSeek V4 makes it impossible to independently verify Hark's 'top-ever' claim
  • Latency measurements were conducted by Hark using its own harness with competing models set to their slowest reasoning levels, limiting the reliability of performance comparisons
  • On WebTailBench v2, one of Hark's own chosen benchmarks, GPT 5.5 outperformed Handoff (72.3 versus 68.6), suggesting the 'best' framing requires qualification
  • The research finding that fewer than 1 in 1000 websites have publicly accessible APIs validates the market need for visual web navigation agents

Monitor whether Hark publishes benchmark results against current-generation models like GPT-5.6 and Opus 5, which have shown substantial gains in computer use tasks. Track real-world deployment outcomes and enterprise adoption rates, particularly for recruiting and customer service use cases. Watch for independent third-party evaluations of Handoff's performance and latency claims using standardized testing conditions.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

DeepSeek Challenges Claude Code with Open Agent Framework
TrendingModel Release

DeepSeek Challenges Claude Code with Open Agent Framework

DeepSeek launched DeepSeek-V4-Pro, an updated flagship model for agentic workloads, alongside DeepSeek Harness v0.1, an open-source agent framework available under MIT license. The releases position DeepSeek as a competitor to Anthropic's Claude Code and OpenAI's Codex by offering developers an alternative agent infrastructure layer. Simultaneously, DeepSeek is shifting from flat API pricing to peak and off-peak rates starting August 16, with substantially higher prices across the board.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI
Google Cuts Gemini Flash Pricing 50% With Faster Iteration
TrendingModel Release

Google Cuts Gemini Flash Pricing 50% With Faster Iteration

Google DeepMind released Gemini 3.7 Flash on August 13, 2026, positioning it as an improved workhorse model for coding and agent-based tasks. The model arrives three weeks after Gemini 3.6 Flash and delivers measurable gains in software engineering, web development, and knowledge-intensive workflows at half the per-token cost of its predecessor. The release reflects developer feedback and algorithmic improvements aimed at production-ready code generation and complex document processing.

· Google Deepmind
Startup Slack Threads Become Commodity for AI Training
TrendingNews

Startup Slack Threads Become Commodity for AI Training

AI training companies like Mercor are actively acquiring internal communications and code from startups, offering payments up to $300,000 for Slack threads, GitHub records, and meeting transcripts. Warmly's CEO received four such acquisition offers within days of the company's HubSpot acquisition announcement. The practice highlights how internal startup data has become a commodity for AI model training, even as acquirers may not want the same datasets.

by Alix Coutures· The Information
Capital One Builds Multi-Agent AI on Customized Open Models
TrendingNews

Capital One Builds Multi-Agent AI on Customized Open Models

Capital One built a multi-agent AI platform centered on customized open-weight models rather than relying on off-the-shelf frontier models. The bank fine-tunes open models with proprietary data and uses a specialized multi-agent orchestration system called MACAW to handle complex workflows like fraud detection and customer service. This approach leverages Capital One's data advantage while enabling extensibility across the enterprise.

· VentureBeat AI