Found first: a primary source the press has not covered yet.
A new paper measures exactly how much the gradient traffic in split learning adds to text reconstruction attacks, and finds it nearly triples the rate of exact document recovery. Georgios Politis and Evangelos Pappas isolate the two signals available to a server-side observer in a split language model setup: the activations sent forward during inference and the gradients sent back during training. The paper is available on arXiv.
What the source says
Experiments ran against GPT-2 and Qwen3-0.6B. On GPT-2, activations alone allow an attacker to recover 94.20% of tokens. Observing gradients in addition raises that to 97.38%, a gain of about 3 percentage points (95% CI: 2.72 to 3.64). The gap is far larger at the document level: without gradients, 13.71% of 32-token documents are reconstructed exactly; with gradients, that figure rises to 37.77%. The strongest defense tested, Secret Mixup, prevents nearly all exact document reconstruction, but an attacker still recovers 83 to 91% of individual tokens.
Why it matters
Split learning exists specifically to keep raw text off a central server, with the premise that sharing activations is a manageable exposure. This paper shows the backward pass carries a separate and substantial attack surface. Any deployment where the server-side operator could observe or log gradient traffic faces a roughly 2.75x higher full-document reconstruction rate than an evaluation limited to the activation threat model. The defense gap compounds the problem: even with Secret Mixup applied, the majority of token content remains recoverable. The authors recommend treating both activations and gradients as sensitive as the raw input text they are derived from.