Procedural Core maps 1M weights to ViT-Base, +2.2pp ImageNet
TL;DR
- Expanding a 1M-parameter recurrent core into an 85M-parameter ViT-Base lifts ImageNet top-1 accuracy by 2.2 pp over random initialization.
- In ViTs the init suppresses high-norm tokens, pushing ImageNet-S mAP from 32.3 to 42.9 and VOC07 CorLoc from 9.9 to 18.4.
- The authors say recurrence is essential for weights that transfer across models, with gains also reported on DINO, FineWeb-Edu, and CodeParrot.
"Transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data." That is the pitch in a new preprint on arxiv by Zachary Shinnick, Christian Internò, Hemanth Saratchandran, Anton van den Hengel and Damien Teney, submitted on 29 September 2026.
The method, called Procedural Core, trains a minimal recurrent transformer on procedurally generated data, then expands its weights to seed transformers of arbitrary width and depth. "Expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization," the authors write.
The downstream numbers run larger than the headline gain. Zero-shot segmentation on ImageNet-S moves from 32.3 to 42.9 mAP. Object localization on VOC07 CorLoc rises from 9.9 to 18.4. Depth estimation on NYUv2 drops RMSE from 1.104 to 0.998. The authors localize the benefit in "the suppression of high-norm tokens" inside the ViT.
Beyond vision, the paper reports improvements on DINO self-supervised training, FineWeb-Edu for natural language, and CodeParrot for code. "Our analysis identifies recurrence as essential for learning compact weights that transfer across models," the paper says. Two researchers we track flagged the link within hours of it going up.
The abstract publishes no training cost for the core itself, and no comparison against supervised pretraining as an initialization baseline.
Shared on Bluesky by 2 AI experts
-
Procedural Core: A Compact Recurrent Initialization for Vision Transformers By @damienteney.bsky.social 's group Idea: train a ViT with parameters tied over layers (="recurrent ViT") on abstract procedurally data. Then …
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Procedural Core: A Compact Recurrent Initialization for Vision Transformers