huggingface.co web signal

Bailo's System Switch Beats Random Deferral, No Doom Exits

Agents ai-business

TL;DR

  • On 900 held-out questions, deferring the least confident 30% of decisions to a reasoning model gained +0.13 [0.08, 0.18] over random deferral.
  • Across 33 closed-loop Doom games and three seeds, no variant of the composed fast-slow agent reached the exit.
  • The rank correlation between the fast actor's AUROC and deferral gain was 0.87, with reasoning carrying about half the gain.

The gate helps on the test set, but the system still cannot find the exit. Gian Luca Bailo's new paper, System Switch, hands control from a fast learned actor to a slow reasoning vision-language model only when a gate opens, while the game keeps running. On 900 held-out questions, deferring the least confident 30 percent of decisions to the reasoner gains +0.13 [0.08, 0.18] over random deferral under one option ordering, and +0.08 [0.02, 0.14] with the options shuffled. 'Reasoning carries about half of it,' the paper reports.

Then the loop closes. Across 33 games and three seeds, 'no variant reaches the exit.'

Bailo serves actor and reasoner through a common llama.cpp interface, with the fast decision models running from 0.15B to 9B parameters. On the error side, zero-shot models 'choose to collect items 1.6-1.8 times more often than chance among their errors,' in any option order, though re-ordering the options shifts some models' accuracy outright.

The rank correlation between the fast actor's AUROC and the gain from deferral is 0.87. The paper frames 'accuracy, calibration and sensitivity' as distinct properties, and reports that the most sensitive model's confidence tracks which kinds of situation it fails, not which answers are wrong.

Thirty-three games. Three seeds. Zero exits.