Paper Finds DPO Beats Representation Steering on LLM Safety
TL;DR
- DPO delivers the strongest overall safety control across a matched evaluation and improves as training data grows.
- Representation steering stays competitive primarily in low-data regimes when the contrastive data is high quality.
- Specialized text monitors lead on detection accuracy; representation probes reach competitive accuracy at substantially lower marginal cost.
DPO delivers the strongest overall safety control for large language models in a matched evaluation, with representation steering holding its own only in narrow conditions, according to a paper posted to arXiv on Sept. 28 by Tianyi Guan, Jianhui Chen and Liangming Pan.
The authors run two tracks. For safety control they compare DPO against three representation steering methods on robustness, practicality and granularity; for monitoring they compare representation probes with fine-tuned and open-weight text monitors on full-response detection, early detection and computational cost. "DPO provides the strongest overall control and generally improves with increasing training data," the abstract states, although its safety "can degrade after subsequent benign fine-tuning." Representation steering "remains competitive primarily in low-data settings, particularly with high-quality contrastive data."
On the monitoring side, "specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost." The paper's most operationally pointed result sits between the two tracks: "monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal." No per-method accuracy figures or specific model names appear in the abstract, and the paper drops into a crowded run of safety coverage on our tracker.
Originally reported by arxiv.org
Read the original article →Original headline: Paper Finds DPO Beats Representation Steering for LLM Safety Except in Low-Data Regimes