Hugging Face Daily Papers · · 5 min read

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A fully open 7B LLM, including model, training infra, data, and wandb. Match Qwen3-8B on general datasets and competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1 on math and agentic search datasets. Its highlight includes FP8 pretraining, MDP midtraining, and AI4AI.</p>\n<p>Code: <a href=\"https://github.com/zgcagi/ZGCM-1\" rel=\"nofollow\">https://github.com/zgcagi/ZGCM-1</a><br>Model: <a href=\"https://huggingface.co/zgcagi/ZGCM-1-7B\">https://huggingface.co/zgcagi/ZGCM-1-7B</a><br>Data: <a href=\"https://huggingface.co/datasets/zgcagi/ZGCM-1-Data\">https://huggingface.co/datasets/zgcagi/ZGCM-1-Data</a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/649aa367c6cf3cc95bc1b7f6/1vahHSXQe7t5mgllVUpxs.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/649aa367c6cf3cc95bc1b7f6/1vahHSXQe7t5mgllVUpxs.png\" alt=\"ZGCM\"></a></p>\n","updatedAt":"2026-09-15T03:29:07.097Z","author":{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","fullname":"Yifei Shen","name":"yshenaw","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6956164240837097},"editors":["yshenaw"],"editorAvatarUrls":["/avatars/4bf5446c261eab08fc06caebf4c5779a.svg"],"reactions":[{"reaction":"🔥","users":["hxkwan","zoeycheang"],"count":2}],"isReport":false}},{"id":"6aa8c1c24d6200af179e4c19","author":{"_id":"657ec96cf010d76b6e42e9ce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/nGpTNDv3myclxQFJLiYZY.jpeg","fullname":"hxkwan","name":"hxkwan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/bnC1HgjoiGEOVXWaFhHhC.png","fullname":"ZGCAGI","name":"zgcagi","type":"org","isHf":false,"details":"Pushing the Frontiers of Intelligence","plan":"team"}},"createdAt":"2026-09-15T03:55:46.000Z","type":"comment","data":{"edited":true,"hidden":false,"latest":{"raw":"A big step toward recursive self-improvement (RSI) 🔥🔥🔥","html":"<p>A big step toward recursive self-improvement (RSI) 🔥🔥🔥</p>\n","updatedAt":"2026-09-15T03:55:55.238Z","author":{"_id":"657ec96cf010d76b6e42e9ce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/nGpTNDv3myclxQFJLiYZY.jpeg","fullname":"hxkwan","name":"hxkwan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/bnC1HgjoiGEOVXWaFhHhC.png","fullname":"ZGCAGI","name":"zgcagi","type":"org","isHf":false,"details":"Pushing the Frontiers of Intelligence","plan":"team"}}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8938113451004028},"editors":["hxkwan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/nGpTNDv3myclxQFJLiYZY.jpeg"],"reactions":[{"reaction":"👍","users":["hxkwan","zoeycheang"],"count":2}],"isReport":false}},{"id":"6aa8cd141956ffeb398f3368","author":{"_id":"68ea189f7b9cd18b10e4e530","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9X1c0SH9ophTKGJWX82qL.png","fullname":"Chuyang Wei","name":"weichy2023","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-15T04:44:04.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"The \"AI4AI\" part is the real headline — agent swarms running cluster ops and data curation means the model partially built itself. RSI era loading… 🚀","html":"<p>The \"AI4AI\" part is the real headline — agent swarms running cluster ops and data curation means the model partially built itself. RSI era loading… 🚀</p>\n","updatedAt":"2026-09-15T04:44:04.833Z","author":{"_id":"68ea189f7b9cd18b10e4e530","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9X1c0SH9ophTKGJWX82qL.png","fullname":"Chuyang Wei","name":"weichy2023","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8947243690490723},"editors":["weichy2023"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9X1c0SH9ophTKGJWX82qL.png"],"reactions":[{"reaction":"🚀","users":["hxkwan"],"count":1}],"isReport":false}},{"id":"6aa8e51da1a609c94366651d","author":{"_id":"674ad1fb7d01cab3e3c0d41e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9KqPYMmvSDaHAR-Sfp3gu.png","fullname":"Yanzhi Zhang","name":"Arthur210","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-15T06:26:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"So cool!!!!! Really surprising work! A big step to RSI&AGI.","html":"<p>So cool!!!!! Really surprising work! A big step to RSI&amp;AGI.</p>\n","updatedAt":"2026-09-15T06:26:37.290Z","author":{"_id":"674ad1fb7d01cab3e3c0d41e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9KqPYMmvSDaHAR-Sfp3gu.png","fullname":"Yanzhi Zhang","name":"Arthur210","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7371867299079895},"editors":["Arthur210"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9KqPYMmvSDaHAR-Sfp3gu.png"],"reactions":[{"reaction":"🤯","users":["hxkwan"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.13356","authors":[{"_id":"6aa8b3615dd4cb9b4cc02579","name":"Jiyan He","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257a","name":"Guang Liang","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257b","name":"Hao Liu","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257c","name":"Haoxiang Guan","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257d","name":"Jinbo Sun","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257e","name":"Junyi Guo","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0257f","name":"Wenjun Feng","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02580","name":"Yantai Xie","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02581","name":"Yifei Shen","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02582","name":"Bin Shao","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02583","name":"Chuyang Wei","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02584","name":"Kai Chen","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02585","name":"Kexin Zhou","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02586","name":"Minghang Zhu","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02587","name":"Shuxin Zheng","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02588","name":"Tie-Yan Liu","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc02589","name":"Taine Zhao","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0258a","name":"Wenhui Zhu","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0258b","name":"Xueyin Xu","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0258c","name":"Xiaoqing Zhang","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0258d","name":"Yatao Li","hidden":false},{"_id":"6aa8b3615dd4cb9b4cc0258e","name":"Yuxuan Ren","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/649aa367c6cf3cc95bc1b7f6/lE6avKx6rG2BI7JR3c_t3.png","https://cdn-uploads.huggingface.co/production/uploads/649aa367c6cf3cc95bc1b7f6/lANH7P4FcbJO4xAxyAMPA.mp4"],"publishedAt":"2026-09-11T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search","submittedOnDailyBy":{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user","name":"yshenaw"},"summary":"In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.","upvotes":125,"discussionId":"6aa8b3615dd4cb9b4cc0258f","projectPage":"https://mp.weixin.qq.com/s/kzScxJki8hY2IHIl32l5cQ","githubRepo":"https://github.com/zgcagi/ZGCM-1","githubRepoAddedBy":"user","ai_summary":"ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.","ai_keywords":["gated sliding-window attention","full attention","FP8 Muon optimizer","Markov Decision Processes","agent swarms","long-context scaling","pre-training","mid-training","post-training","agentic co-training"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":75,"organization":{"_id":"6a97bf85893a0bdea86abc71","name":"zgcagi","fullname":"ZGCAGI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/bnC1HgjoiGEOVXWaFhHhC.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"657ec96cf010d76b6e42e9ce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/nGpTNDv3myclxQFJLiYZY.jpeg","isPro":false,"fullname":"hxkwan","user":"hxkwan","type":"user"},{"_id":"674d191fc406e58320e21dc0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674d191fc406e58320e21dc0/3WUBoDbpGxOu5-OJq0bxS.png","isPro":false,"fullname":"Koutian Wu","user":"ktwu01","type":"user"},{"_id":"62d5ab93c46278c4a9954abc","avatarUrl":"/avatars/67d7148e348416e51cd6af9187756d7a.svg","isPro":false,"fullname":"ustchjy","user":"ustchjy","type":"user"},{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user"},{"_id":"654c3a4009dd7ef52491c080","avatarUrl":"/avatars/c3c14a5e732f7034eb5c50a4a8a47107.svg","isPro":false,"fullname":"Wenjun Feng","user":"USTCKevinF","type":"user"},{"_id":"6a03f485883427d8f45eeacc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/BWYIqDxGMmJ_Rl49OqKyP.jpeg","isPro":false,"fullname":"Zubin Zheng","user":"0SilverBullet","type":"user"},{"_id":"6971eab81bc41d41f2e1c972","avatarUrl":"/avatars/a96a4700d392af05bcc5aa0e2426f8ab.svg","isPro":false,"fullname":"Jiajie Cheng","user":"JaneCheng","type":"user"},{"_id":"664c6083f48f9e269c593105","avatarUrl":"/avatars/8cd69f3811202b6e2a1cb13ee53a8148.svg","isPro":false,"fullname":"Kai Chen","user":"chenkagi","type":"user"},{"_id":"69491ab31b1af1c40a0ac5eb","avatarUrl":"/avatars/241e32efcd52ded0712ec03ab7a6949e.svg","isPro":false,"fullname":"Kailin Deng","user":"iceballoon","type":"user"},{"_id":"6a05825bad1913dfb1f1f0ea","avatarUrl":"/avatars/d33c289e6475b2d0efc6dbee0999f565.svg","isPro":false,"fullname":"Kan Wen","user":"Kanwen","type":"user"},{"_id":"630c43aed73c0a1dcf5c2a78","avatarUrl":"/avatars/1517bb880f1c011e021152eb90aa0f09.svg","isPro":false,"fullname":"tomatobobot","user":"tomatobobot","type":"user"},{"_id":"6aa8bca79bda0f8d805302b0","avatarUrl":"/avatars/90635fcb750088bdf68322c1ecf43d7a.svg","isPro":false,"fullname":"guomeiying","user":"NAITANGGMY","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a97bf85893a0bdea86abc71","name":"zgcagi","fullname":"ZGCAGI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/657ec96cf010d76b6e42e9ce/bnC1HgjoiGEOVXWaFhHhC.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.13356.md","query":{}}">
Papers
arxiv:2609.13356

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Published on Sep 11
· Submitted by
Yifei Shen
on Sep 15
Authors:
,

Abstract

ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.

In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.

Community

Paper submitter about 5 hours ago

A fully open 7B LLM, including model, training infra, data, and wandb. Match Qwen3-8B on general datasets and competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1 on math and agentic search datasets. Its highlight includes FP8 pretraining, MDP midtraining, and AI4AI.

Code: https://github.com/zgcagi/ZGCM-1
Model: https://huggingface.co/zgcagi/ZGCM-1-7B
Data: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data
ZGCM

A big step toward recursive self-improvement (RSI) 🔥🔥🔥

The "AI4AI" part is the real headline — agent swarms running cluster ops and data curation means the model partially built itself. RSI era loading… 🚀

So cool!!!!! Really surprising work! A big step to RSI&AGI.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13356
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.13356 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers