marktechpost.com web signal

Sakana AI's PC-ALM trains 1,000-layer nets without backprop

Sakana AI ai-business

TL;DR

  • Sakana AI's PC-ALM, a layer-local training method, matches backpropagation within about 2 percentage points on a 1,000-layer residual MLP trained on MNIST.
  • On a Fashion-MNIST reference cell of width 32 and depth 32 with ReLU, PC-ALM scores 77.75% versus 78.66% for backprop and 68.13% for plain predictive coding.
  • The MIT-licensed JAX reference implementation runs on CPU and has only been tested on small image benchmarks, including CIFAR-10 and Tiny ImageNet with ResNet-18.

Sakana AI's new predictive coding variant, PC-ALM, trains residual MLPs up to 1,000 layers within about 2 percentage points of backpropagation on MNIST while keeping every weight update local to a single layer. That is the headline claim of the lab's Augmented Lagrangian Predictive Coding paper, released with an MIT-licensed JAX reference implementation on GitHub. As Marktechpost reports, it is "a training method, not a model, and has only been tested on small image benchmarks."

"Backpropagation is a global algorithm: a forward pass, then a backward pass, then a weight update, each locked behind the previous one," the writeup explains. "Brains have no known mechanism for that kind of network-wide phase locking, which is why local-learning alternatives such as predictive coding (PC) keep drawing research interest." Standard PC treats each hidden activation as an optimization variable and penalizes the squared mismatch between an activation and the prediction arriving from the layer below. In deep, narrow networks the supervision signal fades long before it reaches the input, and PC stalls.

PC-ALM attaches a per-layer Lagrange multiplier to that constraint. Inference alternates a primal step on the activations with a dual step that accumulates each layer's prediction error; the team reads the pair as a PI controller, where "the prediction error is the proportional term and the multiplier is the integral term." Zeroing out the dual step recovers ordinary PC.

The numbers are narrow. On a Fashion-MNIST reference cell of width 32, depth 32, ReLU, backprop hits 78.66% test accuracy, PC-ALM 77.75%, and plain PC 68.13%. Gradient cosine similarity to the backprop reference improves from 0.604 to 0.909. The team also runs the method on ResNet-18 with CIFAR-10 and Tiny ImageNet.

What the writeup does not show: any result on transformers, language models, or a benchmark past small vision. The reference code targets CPU.