Linear Aligner Trained in Minutes Boosts TabPFN-3 and TabFM
TL;DR
- A single linear layer, trained on synthetic unlabeled data in seconds to minutes without a GPU, aligns a reduced-context student to a full-context teacher.
- On 38 TabArena classification datasets, the aligned student beat the unaligned baseline in 81% of cases across TabPFN-3 and TabFM.
- At 10% of the training context, aligned models matched unaligned students given two to four times as much data.
The aligner is a single linear layer. It needs no GPU, converges in seconds to minutes on commodity hardware, and sits inside the frozen transformer at the final layer, nudging a data-constrained student's activations toward those of a copy that sees the full training set.
That is the pitch of a new preprint, "Closing the Context Gap: Activation Alignment for Tabular In-Context Learning," posted to arXiv on 6 October. The paper evaluates on 38 classification datasets from the TabArena benchmark, using TabPFN-3 and TabFM, both of which run across 24 in-context-learning transformer blocks. The aligner is fit on synthetic unlabeled queries generated by TabPFN's own unsupervised prior, with 1,000 synthetic samples in the reported runs.
The headline result is modest and specific. "Across both models and all evaluated context budgets, the aligned student outperforms the unaligned baseline in 81% of cases, with statistically significant gains," the paper reports. "With only 10% of the training context, the aligned models achieve predictive parity with unaligned students provided two to four times as much data," and in low-data regimes "alignment recovers nearly half of the teacher's predictive advantage."
Code is posted at github.com/yoel-zeldes/tabalign, one of several recent open-source tabular-ML releases landing this week. The paper does not report wall-clock inference speedups or per-dataset numbers in the abstract.
Originally reported by huggingface.co
Read the original article →Original headline: Activation Alignment Paper Shrinks TabPFN-3 and TabFM Context With a Linear Aligner Trained in Minutes