Recurrent Looped Transformer generalizes parity to 256 bits
TL;DR
- Two RLT splits generalize parity from 40 training bits to 256 bits at 100% accuracy in every seed; an eight-layer Transformer stays at chance.
- On swap-based S_5 permutation tracking at eight times training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer.
- Removing per-token feedback drops parity and S_5 to chance; chunking feedback every four tokens cuts length-64 swap S_5 from 100% to 20%.
Trained on sequences of at most 40 bits, two splits of a new architecture called the Recurrent Looped Transformer generalize parity to 256 bits with 100% accuracy in every seed, according to a preprint from Yifan Zhang, Jichen Feng and Shihan Qin. A matched eight-layer Transformer stays at chance on the same task.
The authors split eight layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so, as the paper puts it, "the computation path grows with sequence length at a fixed per-token cost." They test five splits of the eight layers across three seeds on six algorithmic tasks. On swap-based S_5 permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus "under 1% for the Transformer," and accuracy climbs with decoder depth. On modular arithmetic beyond training lengths, RLT hits 93% against the Transformer's 33%.
The feedback loop is doing the work. Ablations show that removing it "drops parity and swap-based S_5 to chance at every split." Chunked feedback, updating once per four tokens so known tokens inside a chunk can run in parallel, holds 64-bit parity at 99%, but collapses length-64 swap S_5 from 100% to 20%.
Shared on Bluesky by 2 AI experts
-
Recurrent Looped Transformer Zhang et al. Combine context/history transformer with recurrence: have a causal encoder with a recurrent decoder. arxiv.org/abs/2610.07591
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Recurrent Looped Transformer