AdvSim2Real Trains 4B Web Agent Against Adaptive Prompt Injection
TL;DR
- A co-evolved task curriculum, prompt-injection adversary and 4B web agent push attacked completion from 48.07% to 57.48% across 150 tasks.
- Against Kimi-K3, a frontier adversary never seen during training, attacked completion rises 33.6% relative to the base agent.
- Moved to a real browser, strict success climbs from 25.56% to 44.44%, while clean completion rises from 74.89% to 81.33%.
A 4B web agent trained against an adaptive prompt-injection attacker inside a frozen browser simulator raises its task completion under those attacks from 48.07% to 57.48% across 150 web tasks, according to a Hugging Face paper by researchers at Mohamed bin Zayed University of Artificial Intelligence, Amazon and MIT.
The method, AdvSim2Real, co-evolves three things at once rather than fixing any of them. A task curriculum is "rewarded for tasks the agent solves about half of the time", and the adversary is credited, in the authors' phrase, "only for a success flip, an injection that turns a judged success into a failure". It lands in a busy stretch for web-agent safety work on our tracker.
Robustness carries outside training too. Against Kimi-K3, "a frontier-model adversary it never trained against", attacked completion rises 33.6% relative to the base model. When the trained policy is lifted out of the simulator onto a real browser, strict success climbs from 25.56% to 44.44%.
Clean capability does not degrade in the process. It rises alongside robustness, from 74.89% to 81.33% completion on unattacked tasks. The authors frame the result as the agent becoming "both more capable and more robust: its completion rises with and without attacks". Code, the 150-task benchmark and checkpoints are released.
Originally reported by huggingface.co
Read the original article →Original headline: AdvSim2Real Co-Evolves Adversary and Web Agent, Lifts Prompt-Injection Resistance 9 Points