arxiv.org web signal

KaliBench Caps Open Models at 42% on Exact Kali Commands

TL;DR

  • KaliBench scores LLMs on translating analyst intent into exact Kali Linux CLI commands across 8,504 pairs covering 1,642 tools.
  • No open-weight model exceeds 42% exact-command accuracy in the unrestricted setting across 24 tested configurations.
  • SFT and RL with KaliBench's verifiable rewards lift an 8B model to performance comparable with a 685B MoE baseline.

The paper reports that no open-weight language model it tested exceeds 42% exact-command accuracy on Kali Linux CLI tasks in the unrestricted setting. The benchmark, called KaliBench, covers 8,504 query-command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases.

The frame is tight. "Cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag–value bindings, or argument misordering can invalidate execution," the authors write. KaliBench grades on exact commands, not free-form intent.

The 42% ceiling held across 24 configurations of general-purpose and security-focused open-weight models. The paper calls that result a sign of "the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints."

Then the pivot. The authors report that "supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model." The reward signal is runtime-free, built from deterministic canonicalization and alias-aware evaluation rather than live sandbox execution on every grade. A multi-stage verification pipeline combining LLM validation, sandboxed terminal execution, and human-in-the-loop refinement sits behind the dataset itself.

The abstract names no specific models at either end of the 42% band.