NEEDLE paper reports 0% rate on LLM code injection attacks
TL;DR
- NEEDLE is a training-free backdoor removal method that applies weight orthogonalisation after a trigger has been identified.
- The paper reports the lowest mean Attack Success Rate among the defences evaluated, including 0% on code injection attacks.
- The method requires neither a clean reference model nor access to the original poisoned training data.
A new backdoor defence for large language models claims zero attack success rate on code injection, with no retraining, no clean reference model, and no access to the poisoned training data.
The method is called NEEDLE, described in a preprint on arxiv by Minoo Kim, Vasileios Lampos and George Drayson. The authors frame it as "a training-free method for targeted backdoor removal". In their own description of the mechanism, once a trigger has been identified, NEEDLE "estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations."
The headline figure is 0% ASR on "challenging code injection attacks", alongside what the authors call "the lowest mean Attack Success Rate (ASR) among the evaluated defences" and "the lowest KL divergence and minimal changes in capability and safety."
The abstract is thin on scope. It names no specific model families, lists no specific attacks, and does not describe how the trigger gets identified in the first place, a step NEEDLE depends on but does not perform. "Evaluation is conducted across multiple model families and attack types" is all the authors say about the testbed.
Originally reported by paper
Read the original article →Original headline: NEEDLE Removes LLM Backdoors Without Retraining, Hits 0% Attack Rate on Code Injection