Hugging Face Daily Papers · · 4 min read

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

code and data: <a href=\"https://github.com/Gen-Verse/Skill-Entropy-RL\" rel=\"nofollow\">https://github.com/Gen-Verse/Skill-Entropy-RL</a></p>\n","updatedAt":"2026-08-06T02:29:23.155Z","author":{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","fullname":"Ling Yang","name":"Lingaaaaaaa","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7468512058258057},"editors":["Lingaaaaaaa"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05139","authors":[{"_id":"6a73f155c5e410d076869a83","name":"Yinghui He","hidden":false},{"_id":"6a73f155c5e410d076869a84","name":"Ling Yang","hidden":false},{"_id":"6a73f155c5e410d076869a85","name":"Jiarui Liu","hidden":false},{"_id":"6a73f155c5e410d076869a86","name":"Yongjin Yang","hidden":false},{"_id":"6a73f155c5e410d076869a87","name":"Lechen Zhang","hidden":false},{"_id":"6a73f155c5e410d076869a88","name":"Yingcheng Wu","hidden":false},{"_id":"6a73f155c5e410d076869a89","name":"Zhenfei Yin","hidden":false},{"_id":"6a73f155c5e410d076869a8a","name":"Mengdi Wang","hidden":false},{"_id":"6a73f155c5e410d076869a8b","name":"Sanjeev Arora","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning","submittedOnDailyBy":{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","isPro":false,"fullname":"Ling Yang","user":"Lingaaaaaaa","type":"user","name":"Lingaaaaaaa"},"summary":"Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill^2-Bench reveals a skill-switching gap: accuracy decreases on higher-entropy tasks. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill-Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a skill-entropy reward that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill^2-Bench score from 34.4% to 68.4% and from 14.6% to 40.1%, respectively, outperforming competitive baselines. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal. Code available at: https://github.com/Gen-Verse/Skill-Entropy-RL","upvotes":16,"discussionId":"6a73f155c5e410d076869a8c","projectPage":"https://huggingface.co/datasets/Gen-Verse/Skill2-Bench","githubRepo":"https://github.com/Gen-Verse/Skill-Entropy-RL","githubRepoAddedBy":"user","githubStars":3,"organization":{"_id":"64374111a701a7e744c02b0e","name":"princetonu","fullname":"Princeton University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/b3xXusq8Zz3ej8Z6fRTSZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6662a23c2f86097c6d828b96","avatarUrl":"/avatars/2aa31ab30874257529861f2e4024acc2.svg","isPro":false,"fullname":"liu","user":"miao6","type":"user"},{"_id":"69817bd8a819c22bf570fae4","avatarUrl":"/avatars/00642747bfb2bcf443165ae7a4175c0c.svg","isPro":false,"fullname":"Yang","user":"TonyYang1","type":"user"},{"_id":"6662a2ac9ced3e13879c524d","avatarUrl":"/avatars/fa5bb180daad40171c0fde6f5ce081f7.svg","isPro":false,"fullname":"liu","user":"miao66","type":"user"},{"_id":"6981942081c01373225279f5","avatarUrl":"/avatars/b84b75986a3889ccb241b8a25cb09df8.svg","isPro":false,"fullname":"ma","user":"sherryma23","type":"user"},{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","isPro":false,"fullname":"Ling Yang","user":"Lingaaaaaaa","type":"user"},{"_id":"6662a59cf8d1fcc749cbc5de","avatarUrl":"/avatars/0e965b6b996c154b8d39106c0cc5178d.svg","isPro":false,"fullname":"liu","user":"miao99","type":"user"},{"_id":"6981928e9dcf301b73fe1fcb","avatarUrl":"/avatars/42d73f5498c65a4ebe67c168abafa218.svg","isPro":false,"fullname":"yang","user":"martinyang1","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64706ed2fa9fd77212de57ed","avatarUrl":"/avatars/e27564d1ba5bd9ac2f228a9c8e4c9506.svg","isPro":false,"fullname":"Jiarui Liu","user":"Jerry999","type":"user"},{"_id":"698196075cc02c68e5555ea5","avatarUrl":"/avatars/b36383b891f37275386b18f19d0a6ef3.svg","isPro":false,"fullname":"liu","user":"2linda","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6360a4c4c579d1b2c66d4e95","avatarUrl":"/avatars/ba9fee700a4d7d42c073a9dce7d703bc.svg","isPro":false,"fullname":"Lechen Zhang","user":"leczhang","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64374111a701a7e744c02b0e","name":"princetonu","fullname":"Princeton University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/b3xXusq8Zz3ej8Z6fRTSZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05139.md","query":{}}">
Papers
arxiv:2608.05139

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Published on Aug 5
· Submitted by
Ling Yang
on Aug 6
Authors:
,

Abstract

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill^2-Bench reveals a skill-switching gap: accuracy decreases on higher-entropy tasks. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill-Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a skill-entropy reward that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill^2-Bench score from 34.4% to 68.4% and from 14.6% to 40.1%, respectively, outperforming competitive baselines. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal. Code available at: https://github.com/Gen-Verse/Skill-Entropy-RL

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.05139
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.05139 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.05139 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.05139 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers