'Code to Control' writes Python RL policies, skips LLM at runtime
TL;DR
- The method uses an LLM to synthesize a Python controller once, then runs that controller as the policy with no LLM calls at decision time.
- Program structure comes from the LLM, while a derivative-free search fits the controller's parameters by interacting with the environment.
- The authors benchmark on Atari, Flappy Bird, and MuJoCo tasks, claiming the method beats planning-based program synthesis and stays competitive with deep RL.
The paper's offer is odd: use a language model to write a Python controller once, then throw the language model away at runtime. "Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and faster action selection than a PPO policy," authors Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, and Samuel J. Gershman write in the arxiv preprint.
The approach splits the work. The LLM proposes the program structure. A derivative-free search then fits its parameters by interacting with the environment. "Code to Control learns from interaction and reward rather than relying on a task-specific natural-language command and predefined control primitives," the paper states, a jab at a strand of prior work where the LLM is handed a task description and a library of pre-built skills.
Latency is the stated motivation. "Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use," the authors write. Push the LLM's work into code generation and every runtime decision becomes a plain Python function call.
Benchmarks run across Atari, Flappy Bird, and MuJoCo locomotion. The paper claims the method "outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks." The abstract publishes no per-game scores and no interaction counts; the comparisons sit as summary claims pending the full PDF. Two of the researchers we follow posted the paper link.
Shared on Bluesky by 2 AI experts
-
Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman: Code to Control: Synthesizing Parameterized Reactive Controllers https://arxiv.org/abs/2609.38733 https://arxiv.org/pdf/2609.38733 https://arxiv.org/ht…
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Code to Control: Synthesizing Parameterized Reactive Controllers