huggingface.co web signal

Memento 3 Posts 100.0 RHAE on ARC-AGI-3 With Frozen LLM Rulebook

Agents Open Source ai-business

TL;DR

  • Memento 3 clears every level of all 25 public ARC-AGI-3 games with a mean RHAE of 100.0 using 7,518 actions, 44% of the 17,135 human baseline.
  • The agent keeps its underlying LLM frozen and learns by revising a Markdown rulebook plus a compiled Python world-model engine between episodes.
  • On Atari Pong, a learned feedback controller wins three episodes 21:0 after 9,504 emulator frames with no LLM calls during execution.

Memento 3, a system from researchers at University College London and Huawei's Noah's Ark Lab, clears every level of all 25 public ARC-AGI-3 games with a mean RHAE of 100.0, the benchmark's ceiling score, while the underlying language model stays frozen. The paper, posted October 9, pairs a natural-language rulebook stored as a Markdown file with a compiled Python module that the agent rewrites each time its predictions fail.

Authors Haoyu Zhao, Zhengxu Yu and colleagues report the agent used 7,518 total actions across the suite, or 44% of the 17,135-action human baseline. A Claude Opus 5 baseline on the same benchmark scored 40.7 mean RHAE, 59.3 points lower. On a separate Atari Pong case study, a learned feedback controller trained on 9,504 emulator frames wins three episodes 21:0, with no LLM calls during execution. The paper calls this '42x more sample-efficient than model-based RL baselines' such as EfficientZero V2, which it lists at roughly 400,000 frames.

An ablation attributes a 9% action reduction and 18% agent-turn reduction to the rulebook itself. A two-model population run on game wa30 further cuts actions from 899 to 597, a 33% drop while holding RHAE at 100. The result joins a run of agent-memory papers we've tracked on our agents feed over the past week.