arxiv.org web signal

Paper tests the 'cardinal sin' of test-set tuning

TL;DR

  • A new arxiv paper by Fregonara, Viering and van Gemert tests test-set hyperparameter tuning on MNIST-1D, CIFAR-10 and three GLUE tasks.
  • The authors find the inflation effect is real and significant but frequently small relative to other sources of noise.
  • Model rankings remain essentially preserved after test-set tuning, so consistent test tuning may not invalidate benchmarks.

Tuning hyperparameters on the test set, long taught as a 'cardinal sin' of machine learning methodology, may not invalidate benchmarks the way textbooks warn, according to a new arxiv preprint from Matteo Fregonara, Tom Viering and Jan van Gemert.

The authors systematically measure the performance inflation caused by tuning on test data across MNIST-1D, CIFAR-10 and three tasks from the GLUE benchmark. The paper reports that 'while the effect is real and significant, it is frequently small relative to other sources of noise.' In many runs, test-set tuning recovers exactly the same model that validation-set tuning would have selected.

The more consequential claim is about rankings.

The authors write that 'the rankings of models remain essentially preserved after tuning on the test set,' which, if it holds up more broadly, means that consistent test-tuning across competing models may not corrupt leaderboard comparisons.

The paper's benchmarks are modest, so whether the finding generalises to larger modern evaluations is left open. The authors call for 'a more nuanced view' of the practice and ask researchers 'to openly report test tuning' rather than conceal it.

Shared on Bluesky by 2 AI experts