KaliBench Caps Open Models at 42% on Exact Kali Commands
TL;DR
- KaliBench scores LLMs on translating analyst intent into exact Kali Linux CLI commands across 8,504 pairs covering 1,642 tools.
- No open-weight model exceeds 42% exact-command accuracy in the unrestricted setting across 24 tested configurations.
- SFT and RL with KaliBench's verifiable rewards lift an 8B model to performance comparable with a 685B MoE baseline.
The paper reports that no open-weight language model it tested exceeds 42% exact-command accuracy on Kali Linux CLI tasks in the unrestricted setting. The benchmark, called KaliBench, covers 8,504 query-command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases.
The frame is tight. "Cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag–value bindings, or argument misordering can invalidate execution," the authors write. KaliBench grades on exact commands, not free-form intent.
The 42% ceiling held across 24 configurations of general-purpose and security-focused open-weight models. The paper calls that result a sign of "the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints."
Then the pivot. The authors report that "supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model." The reward signal is runtime-free, built from deterministic canonicalization and alias-aware evaluation rather than live sandbox execution on every grade. A multi-stage verification pipeline combining LLM validation, sandboxed terminal execution, and human-in-the-loop refinement sits behind the dataset itself.
The abstract names no specific models at either end of the 42% band.
Originally reported by arxiv.org
Read the original article →Original headline: KaliBench Benchmarks 1,642 Kali Linux Cybersecurity Tools, No Open Model Clears 42% Without Hints