paper web signal

BlindBias jailbreaks black-box LLMs using only sampled text

TL;DR

  • BlindBias reduces jailbreak query cost from 4,000 to 283 API calls on average, a 92.9% drop, with no access to weights or logits.
  • The attack was run against four commercial endpoints: GLM-5, Gemini-3.5-Flash, Qwen3-32B and Kimi-K2.5, across AdvBench, HarmBench and SORRY-Bench.
  • Working code is public on GitHub, and the paper carries no responsible-disclosure or ethics section.

Jailbreaking a frontier commercial LLM through its public API, with no access to weights or token probabilities, now takes an average of 283 queries, down from 4,000. That is the headline cost result in a paper on arXiv dated September 29, 2026, which introduces a method the authors call BlindBias.

The attack targets four commercial endpoints. "We access Gemini-3.5-Flash through Google's native API and GLM-5, Qwen3-32B, and Kimi-K2.5 through the OpenRouter API," the paper states. Earlier controlled-decoding attacks worked by nudging token probabilities, but required the provider to expose numerical logits. "Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text," the authors write. BlindBias reconstructs the next-token distribution from samples alone, then only intervenes at positions it predicts matter.

"Compared with ungated control, the soft pipeline reduces average API calls from 4,000 to 283 (92.9%)," the paper reports.

The authors are modest about priority. They note that "black-box decoding-time attacks already exist when numerical probabilities are exposed," and frame their work narrowly: "Our contribution is to adapt this residual-control mechanism to a stricter, sample-only interface." BlindBias relies on two interface features commonly offered by commercial APIs: repeated sampling from the same prompt, and the ability to supply an assistant prefix for the model to continue from.

Results are reported as "Harm Score" and "Harm Info Score," not the attack-success-rate percentage common elsewhere in the jailbreak literature. On AdvBench's 520 harmful goals, GLM-5 scored highest (Harm Score 4.29), followed by Gemini-3.5-Flash (3.58), Kimi-K2.5 (3.42) and Qwen3-32B (3.03). The authors also evaluate on the 320-behavior test split of HarmBench and the 440 base prompts of SORRY-Bench.

The code is on GitHub at JessonWong/controlled-decoding. The paper contains no responsible-disclosure or ethics section.