techcrunch.com web signal

Hark previews Handoff, its browser agent for real websites

3 sources tracking this story

TL;DR

  • Handoff scored 97.7 on the Online-Mind2Web leaderboard, clearing GPT-5.5 at 92.8, Opus 4.8 at 84.1, and Gemini 2.5 Pro at 69.0 on a human-evaluated third-party benchmark.
  • Handoff's $0.18/$2.37 per million token pricing is less than one-tenth GPT-5.5's $5/$30, making cost a named differentiator alongside benchmark rank.
  • Hark's system predicts the next action rather than the next token, a design departure the company says is fundamental to reliable real-world browser completion.

Brett Adcock's Hark previewed an agent called Handoff that clicks around real websites the way a person does, without official APIs on the other side. TechCrunch reports the agent works on services like Target, Walmart, OpenTable and LinkedIn, analyzing page structure and visual data to decide whether to click a button or type text. The pitch is that the underlying model predicts the next action rather than the next token.

In the demo Adcock ran for TechCrunch, Handoff built a custom flower bouquet from a vague brief that included 'some of the florist's choice.' The company says Handoff is faster and costs significantly less than GPT 5.5 and Opus 4.8, without publishing head-to-head numbers in the preview itself. A parallel Hark press release claims a top score on the industry-standard OM2W benchmark and pricing at less than one-tenth the token price of competing models from Anthropic, OpenAI and Google, with tasks like ordering on DoorDash and Uber Eats and price-comparing flights on United, Delta and American offered as examples.

The context matters. Hark raised $700 million in Series A funding in May 2026 at a $6 billion post-money valuation, and Handoff is still running on a post-trained model, with proper pre-training pushed to later in 2026. That is an unusual sequencing choice, meaning today's demo reflects data pipeline and post-training work, not the full-scale training run Hark still intends to do.

Every claim here is Hark's own, and the TechCrunch preview does not include independent benchmarks or reproducible tests. Neither piece addresses the business question underneath: platforms like LinkedIn and OpenTable have a history of blocking automated agents, and a browser-first agent that mimics a human is exactly the kind of traffic large sites fight. The waitlist is open now, with launch targeted for end of summer 2026, which is when the claims meet real workloads and, likely, real countermeasures.

What others are reporting

Coverage cluster as of 24h after publish

  1. PR Newswire (via Las Vegas Sun) Read →

    Official press release with full benchmark table, pricing breakdown, and the primary Adcock quote framing the product's market rationale.

    The world is full of AI assistants, but you'd never hire an assistant who couldn't use a computer.
  2. Superintelligence News Read →

    Focuses on Hark's action-prediction architecture and its sequenced post-training-before-pre-training strategy; covers real-world browser friction incumbents have not solved.

    The system predicts the next action, not just the next token.