Imprint Reader Turns Weight Updates Into Text, Edits Behavior
TL;DR
- Reader-guided pruning at 0.5% raised harmful-prompt refusal from 57.9% to 64.1% under a safety-maintenance target, per the abstract.
- The same Reader signal drove MetaEdit to lift BFCL Overall from 41.69% to 44.60% without target-task training data.
- On held-out updates the Reader hits judge-based Pass@100 of just 2% for knowledge and 16% for behavior, which the authors frame as feasibility.
At a 0.5% pruning rate guided by a new interpretability tool called the Imprint Reader, harmful-prompt refusal rose from 57.9% to 64.1%, according to a preprint by Guanxu Chen, Qihao Lin and Jing Shao.
The tool is "a model trained with Semantic Mount-and-Read Tuning (SaRT) to describe frozen weight updates" in natural language. The authors frame it as a dual-use signal: it produces text descriptions of what a fine-tune taught a model, and it also "provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update." That proxy is what drives the pruning and editing results.
Beyond the safety number, a related technique the paper calls MetaEdit used behavior descriptions alone, with no target-task training data, and raised BFCL Overall from 41.69% to 44.60%.
The readout itself is still thin. On held-out updates the joint Reader reaches "judge-based Pass@100 of 2% for knowledge and 16% for behavior." The authors describe the work as demonstrating "the feasibility of natural-language readout while pointing to reliability across updates as the next step." It arrives in a busy stretch for safety-adjacent research on our safety tracker.
Originally reported by arxiv.org
Read the original article →Original headline: Imprint Reader Paper Decodes Weight-Update Traces Into Natural Language, Lifts Refusal to 64.1%