techcrunch.com web signal

Inherent Says Faraday Tops Claude, GPT-5.5 at Paper Replication

4 sources tracking this story

TL;DR

  • Faraday at 27B parameters outperformed Claude Opus 4.8 on 73% of Replica tasks in-distribution and maintained its lead on held-out AI-for-science papers.
  • The Replica benchmark spans 310 tasks drawn from 100 ML and AI-for-science papers across NLP, materials science, and weather forecasting, each requiring figure reproduction without the original plot.
  • Long-horizon RL with an auto-generated rubric-based judge, rather than standard fine-tuning or RLHF, is what Inherent credits for embedding scientific intuition into Faraday's weights.

London-based Inherent emerged from stealth with a $50 million seed round and a claim that its "AI scientist" agent, Faraday, beats larger closed models at reproducing published research. The round was led by Index Ventures with Radical Ventures participating, TechCrunch reported.

Faraday is built on a 27-billion-parameter Qwen base and calls OpenAI's GPT-5.5 Codex for coding subtasks. On Inherent's own benchmark, Replica (a set of 310 tasks drawn from 100 machine-learning and AI-for-science papers), the agent outscored Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at reproducing results "without being given the answer in advance," per Inherent.

Co-founder and chief scientist Edward Hughes, a DeepMind alumnus, said the team is "always guided by that north star of building an AI scientist agent." The King's Cross lab is about a dozen people today and plans to reach 20 to 25 by year-end. Co-founder Tantum Collins previously worked on AI policy in the Biden White House.

The stack is the more interesting artifact: a small open-weight "scientist" orchestrating a large closed-weight coder. It is one shape of the pattern our agents tracker has been logging through the summer, as labs mix model tiers rather than default to one frontier system. Cursor this week wired subagent VMs into its cloud agents, a different cut of the same orchestration idea.

Replica is Inherent's own benchmark. No independent rerun of the Claude Opus 4.8 or GPT-5.5 comparisons has been published.

What others are reporting

Coverage cluster as of 24h after publish

  1. Primary technical source: 47 pages, 12 figures, detailing Replica benchmark design, the rubric-based judge methodology, and full long-horizon RL training approach behind Faraday.

    Faraday outperformed Claude on 73% of tasks in-distribution and maintained a lead on held-out AI-for-science papers.
  2. Tech.eu Read →

    Earliest outlet on the stealth exit; adds Matt Clifford's endorsement and frames Inherent as reimagining the scientific method rather than layering AI onto existing research workflows.

    A system designed to help humans and self-improving AI work together on genuine scientific discovery.
  3. The Next Web Read →

    Focuses on Inherent's public benefit corporation structure and Tantum Collins' White House AI policy background, both absent from TechCrunch's benchmark-led coverage.

    Most AI is built to answer questions. What it can't do yet is figure out which questions are worth asking.