Code: <a href=\"https://github.com/AI9Stars/AStar-Thought\" rel=\"nofollow\">https://github.com/AI9Stars/AStar-Thought</a></p>\n","updatedAt":"2026-09-09T09:20:16.753Z","author":{"_id":"654112124a48289000a34245","avatarUrl":"/avatars/f02420326f97c7ad1fc6350b109791b8.svg","fullname":"Xu Xiaoang","name":"xxang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.811453104019165},"editors":["xxang"],"editorAvatarUrls":["/avatars/f02420326f97c7ad1fc6350b109791b8.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.07821","authors":[{"_id":"6aa12470ff4bf7311191aa8b","user":{"_id":"654112124a48289000a34245","avatarUrl":"/avatars/f02420326f97c7ad1fc6350b109791b8.svg","isPro":false,"fullname":"Xu Xiaoang","user":"xxang","type":"user","name":"xxang"},"name":"Xiaoang Xu","status":"claimed_verified","statusLastChangedAt":"2026-09-09T09:22:39.827Z","hidden":false},{"_id":"6aa12470ff4bf7311191aa8c","name":"Siyuan Liu","hidden":false},{"_id":"6aa12470ff4bf7311191aa8d","name":"Shuo Wang","hidden":false},{"_id":"6aa12470ff4bf7311191aa8e","name":"Junlan Feng","hidden":false},{"_id":"6aa12470ff4bf7311191aa8f","name":"Fanyu Meng","hidden":false},{"_id":"6aa12470ff4bf7311191aa90","name":"Zhu Zhang","hidden":false},{"_id":"6aa12470ff4bf7311191aa91","name":"Jixun Wang","hidden":false},{"_id":"6aa12470ff4bf7311191aa92","name":"Xiaorong Wang","hidden":false},{"_id":"6aa12470ff4bf7311191aa93","name":"Zihan Zhou","hidden":false},{"_id":"6aa12470ff4bf7311191aa94","name":"Xin Li","hidden":false},{"_id":"6aa12470ff4bf7311191aa95","name":"Chaojun Xiao","hidden":false},{"_id":"6aa12470ff4bf7311191aa96","name":"Yiming Zhang","hidden":false},{"_id":"6aa12470ff4bf7311191aa97","name":"Huijia Wu","hidden":false},{"_id":"6aa12470ff4bf7311191aa98","name":"Liuyu Xiang","hidden":false},{"_id":"6aa12470ff4bf7311191aa99","name":"Peipei Li","hidden":false},{"_id":"6aa12470ff4bf7311191aa9a","name":"Zhaofeng He","hidden":false}],"publishedAt":"2026-09-07T17:56:20.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM","submittedOnDailyBy":{"_id":"654112124a48289000a34245","avatarUrl":"/avatars/f02420326f97c7ad1fc6350b109791b8.svg","isPro":false,"fullname":"Xu Xiaoang","user":"xxang","type":"user","name":"xxang"},"summary":"Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29times, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.","upvotes":9,"discussionId":"6aa12470ff4bf7311191aa9b","ai_summary":"A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories to selectively retain explicit reasoning steps or compress them into continuous latent tokens, improving accuracy and efficiency.","ai_keywords":["A*-Thought-V2","chain-of-thought","hidden-state trajectory","explicit-implicit interleaved latent architecture","3D PCA space","directional angles","stepwise embedding forcing","label forcing","soft multi-modal vocabulary distribution","latent tokens"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"654112124a48289000a34245","avatarUrl":"/avatars/f02420326f97c7ad1fc6350b109791b8.svg","isPro":false,"fullname":"Xu Xiaoang","user":"xxang","type":"user"},{"_id":"6a6c84f892fd458eaa9c7938","avatarUrl":"/avatars/6390b9321a236bf70ff45cf3b85a0ccf.svg","isPro":false,"fullname":"Jennifer Williams","user":"velvetscope","type":"user"},{"_id":"6a6c8a1109f0af0927dee1c0","avatarUrl":"/avatars/f6e88ed87230495e6a3bda2d6cbd8aef.svg","isPro":false,"fullname":"James Anderson","user":"Harbor-James","type":"user"},{"_id":"6a6deafdef16fe7cceb88c8f","avatarUrl":"/avatars/a93a19a0d494f81ec71de4742af03809.svg","isPro":false,"fullname":"Charles Harris","user":"lunarflow","type":"user"},{"_id":"6a8254d04030eff1e05f6051","avatarUrl":"/avatars/2ee14b86c51ced545c4064649e39f5d5.svg","isPro":false,"fullname":"James Garcia","user":"EmberTrail78","type":"user"},{"_id":"6a9b52c79379d331be8ee7c7","avatarUrl":"/avatars/7260a066d244d3f499f48cd96a9823b1.svg","isPro":false,"fullname":"손성진","user":"seongjin5777","type":"user"},{"_id":"6aa0efc38efdd8def0ccf0d3","avatarUrl":"/avatars/014466eb843af341005931ff1f18e063.svg","isPro":false,"fullname":"Anthony Sanchez","user":"Velvet-Edward16","type":"user"},{"_id":"6615494716917dfdc645c44e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6615494716917dfdc645c44e/M_aMgH5QzSl_hm-srhJwW.webp","isPro":true,"fullname":"Daniel B.","user":"FlameF0X","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.07821.md","query":{}}">
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
Abstract
A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories to selectively retain explicit reasoning steps or compress them into continuous latent tokens, improving accuracy and efficiency.
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29times, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.07821 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.07821 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.07821 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.