Hugging Face Daily Papers · · 4 min read

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A large-scale new architecture with latent space prediction.</p>\n","updatedAt":"2026-09-11T02:52:17.518Z","author":{"_id":"6529f79e802e3d1a4f8ec662","avatarUrl":"/avatars/d05320c370a6497d8792ef5acb563dd5.svg","fullname":"Yuliang Liu","name":"yuliang03181","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8301075100898743},"editors":["yuliang03181"],"editorAvatarUrls":["/avatars/d05320c370a6497d8792ef5acb563dd5.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10715","authors":[{"_id":"6aa36c8547a406da7901e71d","name":"NCP Team","hidden":false},{"_id":"6aa36c8547a406da7901e71e","user":{"_id":"6507c84ec53e1a7f17d996fd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/U2qzEGvhPk7RjINeaIYid.jpeg","isPro":false,"fullname":"Maximus Cao","user":"Clover-Hill","type":"user","name":"Clover-Hill"},"name":"Jiaqi Cao","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:53:24.093Z","hidden":false},{"_id":"6aa36c8547a406da7901e71f","name":"Chiyu Chen","hidden":false},{"_id":"6aa36c8547a406da7901e720","name":"Shuang Cheng","hidden":false},{"_id":"6aa36c8547a406da7901e721","name":"Xu Cheng","hidden":false},{"_id":"6aa36c8547a406da7901e722","name":"Beiya Dai","hidden":false},{"_id":"6aa36c8547a406da7901e723","name":"Yufan Feng","hidden":false},{"_id":"6aa36c8547a406da7901e724","name":"Kewen Ge","hidden":false},{"_id":"6aa36c8547a406da7901e725","name":"Ruijun Ge","hidden":false},{"_id":"6aa36c8547a406da7901e726","name":"Jiayi Huang","hidden":false},{"_id":"6aa36c8547a406da7901e727","name":"Yang Jiao","hidden":false},{"_id":"6aa36c8547a406da7901e728","name":"Dahua Lin","hidden":false},{"_id":"6aa36c8547a406da7901e729","name":"Zhouhan Lin","hidden":false},{"_id":"6aa36c8547a406da7901e72a","name":"Yifan Liu","hidden":false},{"_id":"6aa36c8547a406da7901e72b","name":"Yuliang Liu","hidden":false},{"_id":"6aa36c8547a406da7901e72c","name":"Biqing Qi","hidden":false},{"_id":"6aa36c8547a406da7901e72d","name":"Mowen Ruan","hidden":false},{"_id":"6aa36c8547a406da7901e72e","name":"Junzhe Shen","hidden":false},{"_id":"6aa36c8547a406da7901e72f","name":"Yunchong Song","hidden":false},{"_id":"6aa36c8547a406da7901e730","name":"Hao Sun","hidden":false},{"_id":"6aa36c8547a406da7901e731","name":"Zhongbo Tian","hidden":false},{"_id":"6aa36c8547a406da7901e732","name":"Yixuan Wang","hidden":false},{"_id":"6aa36c8547a406da7901e733","user":{"_id":"6609a53bd81d611249ef5266","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6609a53bd81d611249ef5266/h31hdQFl-jhRnO8R6Gr4C.png","isPro":false,"fullname":"Rubin Wei","user":"Rubin-Wei","type":"user","name":"Rubin-Wei"},"name":"Rubin Wei","status":"claimed_verified","statusLastChangedAt":"2026-09-11T08:45:04.740Z","hidden":false},{"_id":"6aa36c8547a406da7901e734","name":"Jiaxin Xiong","hidden":false},{"_id":"6aa36c8547a406da7901e735","name":"Kangyu Yang","hidden":false},{"_id":"6aa36c8547a406da7901e736","name":"Qian Yao","hidden":false},{"_id":"6aa36c8547a406da7901e737","name":"Qi Zhang","hidden":false},{"_id":"6aa36c8547a406da7901e738","name":"Bowen Zhou","hidden":false}],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction","submittedOnDailyBy":{"_id":"6529f79e802e3d1a4f8ec662","avatarUrl":"/avatars/d05320c370a6497d8792ef5acb563dd5.svg","isPro":false,"fullname":"Yuliang Liu","user":"yuliang03181","type":"user","name":"yuliang03181"},"summary":"We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.","upvotes":122,"discussionId":"6aa36c8647a406da7901e739","ai_summary":"NCP-ArchPreview is a large latent-space language model that jointly trains next-token and next-concept prediction to improve pretraining efficiency and downstream performance.","ai_keywords":["latent-space language model","next-token prediction","Next Concept Prediction","product-quantized concept vocabulary","Concept Module","autoregressive generation","domain adaptation","DFlash2 drafter"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6529f79e802e3d1a4f8ec662","avatarUrl":"/avatars/d05320c370a6497d8792ef5acb563dd5.svg","isPro":false,"fullname":"Yuliang Liu","user":"yuliang03181","type":"user"},{"_id":"6a5444dca6be37d0b048215f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a5444dca6be37d0b048215f/tN_FXokznw7tFzlB6Ca35.png","isPro":false,"fullname":"Jiarui Wang","user":"Jiarui-Wang","type":"user"},{"_id":"68b013324640ce38c97de573","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/At5XpAwgszfYXaXRtJIJ6.png","isPro":false,"fullname":"Yifan Liu","user":"PaulGoodman0700","type":"user"},{"_id":"66d7ffaeccb4a994e3a8f209","avatarUrl":"/avatars/53b102ef2ac32c69b28931bcce10b6db.svg","isPro":false,"fullname":"sam","user":"songyunchong","type":"user"},{"_id":"6530c9d7d107f378e105d667","avatarUrl":"/avatars/889dfcb6514c90351802bebb4a34a78f.svg","isPro":false,"fullname":"Junzhe Shen","user":"JunzheS","type":"user"},{"_id":"66458107219ad12f47bc8fd4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66458107219ad12f47bc8fd4/8NqMRPB2Ko4GIOtL7ZzOj.jpeg","isPro":false,"fullname":"Yixuan Wang","user":"LuckyOrz","type":"user"},{"_id":"6943c6fd3c7aac2b51c1963f","avatarUrl":"/avatars/57a640461755a2dd1145d92af33447fd.svg","isPro":false,"fullname":"Kangyu Yang","user":"KangyuYang","type":"user"},{"_id":"68f6e9ebbbb17a372ec66733","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/hzPbVVWD8SmmWCMJ-jr9r.png","isPro":false,"fullname":"Yixuan Wang (SII)","user":"SII-Lucky","type":"user"},{"_id":"69855c4acdc038b0a77f1514","avatarUrl":"/avatars/1ac66abc06d2803b0a2a344acd980ce7.svg","isPro":false,"fullname":"Hao Doou","user":"Hao126","type":"user"},{"_id":"6925b9449cb916de56ccf8fe","avatarUrl":"/avatars/41d242dc0770536e7e11acf2c1994828.svg","isPro":false,"fullname":"Qingyu Shi","user":"xiawuuuliam","type":"user"},{"_id":"673f15217d6de164a4f53f43","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FAiaairwZuQSNkGUGRQqP.png","isPro":false,"fullname":"Yixuan Wang","user":"LuckyXH","type":"user"},{"_id":"69398b1fea4ada619c51a4df","avatarUrl":"/avatars/365e9ef6f897dd945987d06c56112abc.svg","isPro":false,"fullname":"熊佳鑫","user":"DichenXiong","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"query":{}}">
Papers
arxiv:2609.10715

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

Published on Sep 9
· Submitted by
Yuliang Liu
on Sep 11
#1 Paper of the day
Authors:
,

Abstract

NCP-ArchPreview is a large latent-space language model that jointly trains next-token and next-concept prediction to improve pretraining efficiency and downstream performance.

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.

Community

A large-scale new architecture with latent space prediction.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Browse 18 models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.10715 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.10715 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers