Seven-Person Team's Cyber Agent Hits 63.24% on CyberGym, Ranks 1st at Comparable Scale

Found first: a primary source the press has not covered yet.

A seven-person independent team submitted Feyospace-v1 on September 8, 2026, showing that their open-weight cyber agents rank first among all models at comparable parameter scales on the CyberGym benchmark. The top checkpoint, Feyospace-s1, records a 63.24% verified success rate, placing 10th overall on the leaderboard as of September 1, 2026.

What the source says

The paper refers to the team as the Cyber Mercury Seven, with Zongjie Li and six co-authors. They produced 164,269 long-context supervised fine-tuning trajectories using a five-component data engine: Choulea (hidden-reasoning analysis), SkyReal (reduced teacher-sampling cost), Hongzwang (API restriction bypass), PSBreakup (capability restoration after model merging), and Kreator (converting expert interventions into trainable reasoning). Training environments span coding, vulnerability research, CTF challenges, kernel history, full-exploit scenarios, firmware, and device-backed tasks. Average gains over baseline are 23.76% on the full CyberGym suite and 10.49% across pooled CTF suites. All three released checkpoints rank first among models at comparable parameter scales.

Why it matters

The result puts a concrete number on how far offensive cyber AI capability has moved outside major labs. A team of seven, operating independently, placed a model in the top ten of a major agentic cybersecurity benchmark and outranked every model at a comparable parameter scale. The pipeline is described in enough detail to be replicated. The checkpoints are released, meaning they can be downloaded and further fine-tuned.