Hugging Face Daily Papers · · 5 min read

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We present an end-to-end, data-centric framework for post-training open-weight language models as agentic cyber systems. The framework addresses five practical challenges in capability transfer: reasoning-signature analysis with Choulea(\"思维链蒸馏\"), cost-efficient teacher sampling with SkyReal(“中转杠杆”), constrained teacher elicitation with Hongzwang(“无限制越狱”), recovery of capabilities weakened by model merging through PSBreakup(“模型拆版”), and the conversion of expert interventions into trainable reasoning with Kreator(“专家融合”).</p>\n<p>We construct resettable, environment-grounded tasks spanning repository-level coding, vulnerability reproduction, CTFs, Linux kernel history, exploit development, firmware, and device-backed systems. Candidate trajectories are retained only after execution verification and evidence auditing, resulting in 164,269 audited trajectories for long-context supervised fine-tuning.<br>Without reinforcement learning, the resulting Feyospace checkpoints achieve an average improvement of 23.76% on CyberGym and 10.49% across pooled CTF suites. In the evaluation snapshot reported in the paper, Feyospace-s1 reaches a 63.24% verified success rate on CyberGym and ranks 10th overall, while all three checkpoints rank first among open models at comparable parameter scales.</p>\n","updatedAt":"2026-09-14T06:52:57.545Z","author":{"_id":"6aa3d6b6f5ca096ba4db482f","avatarUrl":"/avatars/49311dca3dbc3cc65540ed47375fe461.svg","fullname":"zongjie li","name":"zongkey","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8734341263771057},"editors":["zongkey"],"editorAvatarUrls":["/avatars/49311dca3dbc3cc65540ed47375fe461.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08418","authors":[{"_id":"6aa12e34ff4bf7311191aab8","user":{"_id":"6aa3d6b6f5ca096ba4db482f","avatarUrl":"/avatars/49311dca3dbc3cc65540ed47375fe461.svg","isPro":false,"fullname":"zongjie li","user":"zongkey","type":"user","name":"zongkey"},"name":"Zongjie Li","status":"claimed_verified","statusLastChangedAt":"2026-09-11T12:32:37.187Z","hidden":false},{"_id":"6aa12e34ff4bf7311191aab9","name":"Alan Z. W","hidden":false},{"_id":"6aa12e34ff4bf7311191aaba","name":"John Nicolas J","hidden":false},{"_id":"6aa12e34ff4bf7311191aabb","name":"Walter H. F","hidden":false},{"_id":"6aa12e34ff4bf7311191aabc","name":"Scott Donald L","hidden":false},{"_id":"6aa12e34ff4bf7311191aabd","name":"Gordon Y. P","hidden":false},{"_id":"6aa12e34ff4bf7311191aabe","name":"Deke X Jr","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models","submittedOnDailyBy":{"_id":"6aa3d6b6f5ca096ba4db482f","avatarUrl":"/avatars/49311dca3dbc3cc65540ed47375fe461.svg","isPro":false,"fullname":"zongjie li","user":"zongkey","type":"user","name":"zongkey"},"summary":"Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.","upvotes":14,"discussionId":"6aa12e34ff4bf7311191aabf","projectPage":"https://verapraxis.ai/feyospace","ai_summary":"A data-centric framework with specialized systems for reasoning analysis, cost reduction, and execution verification enables small teams to train open-weight cyber agents that achieve top-tier performance on benchmark suites.","ai_keywords":["post-training","supervised fine-tuning","model merging","reasoning signatures","execution verification","evidence auditing","long-context supervised fine-tuning","agentic cyber capability"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6a982d49dd257230cc8e310d","name":"feyospace","fullname":"feyospace","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/696761f658bab21869b21d9b/NHLMkll649mzxmPkcyXtE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"696761f658bab21869b21d9b","avatarUrl":"/avatars/cd4b24260da94a5e5b81b1c2ad7ed936.svg","isPro":false,"fullname":"zongjie li","user":"feyospacehacker","type":"user"},{"_id":"687620b15a6752969a51f75f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/QkNF4z2C_IBM9gxQ1XED8.png","isPro":false,"fullname":"Sprout Huang","user":"HRXUST","type":"user"},{"_id":"64ba5484db1774c08e1a531c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba5484db1774c08e1a531c/vZfJ7y5w7XFSlEr5RccLn.jpeg","isPro":false,"fullname":"Donald Lam","user":"mhlam","type":"user"},{"_id":"6aa13d790f66b7f94645453f","avatarUrl":"/avatars/bc468174ecc80bb9b3d22e930f6a5b5d.svg","isPro":false,"fullname":"ZebW","user":"ZebBlockchain","type":"user"},{"_id":"67c7f5c6ec8fcf71f395dbf9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/LJhTn8WKCDXPSxBpZS7U9.png","isPro":false,"fullname":"Dai","user":"MarcusD7","type":"user"},{"_id":"621731ee29500f41901123d5","avatarUrl":"/avatars/ec413c33eb77b67dba35e03c4857f222.svg","isPro":false,"fullname":"YUANHANG YANG","user":"ysngkil","type":"user"},{"_id":"6aa14e36d87071cdcee37a53","avatarUrl":"/avatars/dea0dcaf53c5e6014b101eec3344a6af.svg","isPro":false,"fullname":"aha","user":"ahamountain","type":"user"},{"_id":"6400b125a3b8fe3ac0ecb6f3","avatarUrl":"/avatars/9a17f6064a71a944e770860239688654.svg","isPro":false,"fullname":"Ning Lu","user":"ColinLu50","type":"user"},{"_id":"6354cb2f74026d8c3afe9adc","avatarUrl":"/avatars/2b14f9de396ae3dadb6be569f81dcd1e.svg","isPro":false,"fullname":"zongjieli","user":"zongjieli","type":"user"},{"_id":"6aa3d6b6f5ca096ba4db482f","avatarUrl":"/avatars/49311dca3dbc3cc65540ed47375fe461.svg","isPro":false,"fullname":"zongjie li","user":"zongkey","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a982d49dd257230cc8e310d","name":"feyospace","fullname":"feyospace","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/696761f658bab21869b21d9b/NHLMkll649mzxmPkcyXtE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08418.md","query":{}}">
Papers
arxiv:2609.08418

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

Published on Sep 8
· Submitted by
zongjie li
on Sep 14
Authors:

Abstract

A data-centric framework with specialized systems for reasoning analysis, cost reduction, and execution verification enables small teams to train open-weight cyber agents that achieve top-tier performance on benchmark suites.

Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.

Community

Paper author Paper submitter about 3 hours ago

We present an end-to-end, data-centric framework for post-training open-weight language models as agentic cyber systems. The framework addresses five practical challenges in capability transfer: reasoning-signature analysis with Choulea("思维链蒸馏"), cost-efficient teacher sampling with SkyReal(“中转杠杆”), constrained teacher elicitation with Hongzwang(“无限制越狱”), recovery of capabilities weakened by model merging through PSBreakup(“模型拆版”), and the conversion of expert interventions into trainable reasoning with Kreator(“专家融合”).

We construct resettable, environment-grounded tasks spanning repository-level coding, vulnerability reproduction, CTFs, Linux kernel history, exploit development, firmware, and device-backed systems. Candidate trajectories are retained only after execution verification and evidence auditing, resulting in 164,269 audited trajectories for long-context supervised fine-tuning.
Without reinforcement learning, the resulting Feyospace checkpoints achieve an average improvement of 23.76% on CyberGym and 10.49% across pooled CTF suites. In the evaluation snapshot reported in the paper, Feyospace-s1 reaches a 63.24% verified success rate on CyberGym and ranks 10th overall, while all three checkpoints rank first among open models at comparable parameter scales.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.08418
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.08418 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.08418 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.08418 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers