huggingface.co web signal

Prefill Cues Match RL Math Gains in Olmo-3-7B and Qwen3-14B

TL;DR

  • A `.\n\n Okay` prefill lifts Olmo-3-7B's MATH-500 pass@1 from 42% to 78%, matching the 75% its RL-trained counterpart achieves.
  • KL divergence between base and RL models peaks at the first two output tokens, suggesting RL mostly concentrates probability on cues already in pretraining.
  • Causal data edits turn nonsense cues into working ones: `.\n\n Chicken` takes Olmo-3-7B's MATH-500 from 2.4% to 37.2% after a corpus swap.

Prefilling a short cue at the start of a base model's response can push its math accuracy to roughly match a reinforcement-learned counterpart. In a paper posted on Hugging Face, the authors report that forcing Olmo-3-7B to begin with `.\n\n Okay` raises its MATH-500 pass@1 from 42% to 78%, with the model's RL-trained counterpart landing at 75%. On Qwen3-14B, the cue `␣Alright,` lifts the same benchmark from 72% to 87%.

The abstract frames the result as evidence that RL is amplifying patterns already in pretraining rather than installing new reasoning: "RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model." KL divergence between the base and RL models peaks on the first two output tokens, and the probability of `.\n\n Okay` in Olmo-3-7B rises from 0.14 to 0.65 after RL. For Qwen3-14B, the `␣Alright,` cue climbs from 0.04 to 0.58.

The causal data intervention is the sharper piece. Swapping 'okay' for 'chicken' in Olmo-3-7B's mid-training corpus turns a nonsense cue into a working one, with `.\n\n Chicken` taking MATH-500 accuracy from 2.4% to 37.2% and GSM8K from 2.2% to 60.6%. A parallel edit, the paper reports, makes the prompt instruction 'Think duck duck goose' as effective as 'Think step by step' at eliciting reasoning.

Different cues also pull hidden-state representations toward different document types in the training set: `.\n\n Okay` toward synthetic reasoning traces, `To` toward expository math, `Answer` toward short Q-and-A documents. Checking phrases appear in 97% of `Okay`-cued responses and 5% of `To`-cued ones.

A safety case study runs the same framework. The cue `␣I'm␣sorry` biases models toward refusing both benign and harmful prompts, while `␣Okay,` lowers refusal and lifts the harmful-response rate on unsafe requests to 42.4%. The result lands in a run of recent RL-vs-pretraining writeups tracked on AI Weekly's fine-tuning page, each probing what RL post-training really changes beyond the base model.