paper web signal

Feyospace-v1 hits 63.24% on CyberGym in seven-author paper

TL;DR

  • Feyospace-s1 posts a 63.24% verified success rate and ranks 10th on the official CyberGym leaderboard as of September 1, 2026.
  • A seven-person independent team trained three checkpoints, each ranking first among open-weight models at comparable parameter scale.
  • The data pipeline yielded 164,269 execution-verified trajectories and lifted CyberGym scores by 23.76% on average across checkpoints.

A seven-person independent team says it fine-tuned open-weight models that now rank first at their parameter scale on CyberGym, an agentic offensive-security benchmark. The arXiv paper, submitted September 8, 2026, describes Feyospace-v1, whose flagship checkpoint Feyospace-s1 posts a "verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard" as of September 1, 2026.

The framing is deliberate. "Training capable cyber agents is often treated primarily as a problem of model scale," the abstract states, "yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers." The team's answer is not a larger model but a data pipeline.

That pipeline has five named parts. Choulea analyzes hidden reasoning signatures. SkyReal reduces teacher-sampling cost. Hongzwang "bypasses API restrictions on teacher execution." PSBreakup restores capabilities weakened by model merging. Kreator converts expert interventions into trainable reasoning. Fed through resettable environments spanning coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed tasks, the process yielded 164,269 trajectories retained only after execution verification and evidence auditing.

Across three checkpoints, gains average 23.76% on the full CyberGym suite and 10.49% across pooled CTF suites. The paper's own summary is unusually explicit: "To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability."

Neither the base models the team started from, the compute budget, nor any release plan for the weights appears in the abstract.

Shared on Bluesky by 1 AI expert