arxiv.org web signal

Kevin Murphy unifies two hierarchical RL frameworks via HHMM

TL;DR

  • Kevin Murphy's arxiv note, filed September 13, 2026, unifies two recent proposals for goal-based hierarchical reinforcement learning.
  • The note merges Tasse et al.'s agent-centric general value function with Murphy's 2025 belief-state agent, using hierarchical hidden Markov models.
  • The ACGVF framework lets agents choose both which goal to pursue and when to declare it finished, but assumes full observability.

Kevin Murphy's new arxiv note, filed September 13, 2026, is plumbing: it unifies two recent proposals for goal-based hierarchical reinforcement learning under the formalism of hierarchical hidden Markov models.

The merge target is Tasse et al.'s agent-centric general value function construction, which Murphy describes as letting the agent "make two decisions that are normally imposed by the environment or agent designer: which goal to pursue and when to declare a goal as finished." The catch, he writes, is that the ACGVF "assumes the environment is fully observed." His own 2025 proposal handled partial observability through an internal belief state z_t, but "the goals were assumed to be externally provided." This note stitches the two together.

No experiments. No benchmarks. The abstract reports a formalism and nothing more.

Shared on Bluesky by 1 AI expert