Found first: a primary source the press has not covered yet.
Researchers from Adobe, Amazon, and the University of Southern California have published a controlled decoding attack on black-box LLMs that requires no access to model weights or token probabilities, only sampled text output. The method, BlindBias, achieves the highest mean harm score in 20 of 24 benchmark comparisons against existing jailbreak techniques.
What the source says
The paper, by Jesson Wang, Shawn Li, Wei Yang, and Franck Dernoncourt (Adobe), Ryan A. Rossi (Amazon), and Charith Peris and Yue Zhao (University of Southern California), introduces BlindBias. The method reconstructs token probability distributions from sampled outputs without needing logits, then concentrates its control signal at the small subset of positions where distributions shift most along successful attack trajectories. A speculative prefix-verification step cuts required API calls from 4,000 to 283, a 92.9% reduction over the ungated baseline. Tested against GLM-5, Gemini-3.5-Flash, Qwen3-32B, and Kimi-K2.5 on AdvBench, HarmBench, and SORRY-Bench, BlindBias records the highest mean harm score in 20 of 24 comparisons against PAIR, GPTFuzz, LogiBreak, and FlipAttack. Peak scores are 4.29 out of 5 on GLM-5/AdvBench and 3.67 out of 5 on Gemini-3.5-Flash/SORRY-Bench.
Why it matters
Prior controlled decoding attacks required access to model weights or token probabilities. BlindBias works without either, making the technique applicable to any inference endpoint that returns text. The 92.9% reduction in API calls also lowers the operational cost of mounting attacks at scale. Safety teams at labs and enterprises running inference APIs now face a class of attack designed specifically for the interfaces they expose.