Split-Learning Gradients Nearly Triple Document Leakage
TL;DR
- On GPT-2, exact 32-token document reconstruction rose from 13.71% without gradients to 37.77% when the attacker also watched gradient traffic.
- Per-token recovery climbed from 94.20% to 97.38% with gradients included, a 3.17-point gain (95% CI [2.72, 3.64]).
- The tested 'secret mixup' defense nearly eliminated exact document reconstruction but still left 83-91% of individual tokens recoverable.
On GPT-2, an observer sitting at the split between client and server rebuilt the client's exact 32-token documents 37.77% of the time when it could watch gradient traffic. From activations alone, the same attack succeeded 13.71% of the time.
The figures come from a preprint on arxiv by Georgios Politis and Evangelos Pappas, measuring what the backward pass adds to text leakage in split language models. Per-token recovery climbed from 94.20% to 97.38% with gradients included, a 3.17-point gain with a 95% confidence interval of [2.72, 3.64].
"We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help," the authors write in the abstract.
They evaluate one defense, 'secret mixup', which blends outgoing vectors with decoys. It nearly eliminated exact document reconstruction. Per-token recovery still ran 83 to 91 percent.
The experiments cover GPT-2 and Qwen3-0.6B and track how the chosen split layer trades model quality against information leakage. The authors' recommendation to practitioners: report leakage both per-token and per-document, and treat split-model outputs with security sensitivity equivalent to raw text.
Originally reported by paper
Read the original article →Original headline: Split-LLM Training Gradients Boost Full-Document Reconstruction 2.7x Beyond Activations Alone