Hugging Face Daily Papers · · 5 min read

ISO: An RLVR-Native Optimization Stack

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🪐 ISO: Inherit the spectrum, optimize the frames.</p>\n<p><strong>The finding.</strong> RLVR barely rewrites the base model's singular-value spectra, and the new capability lives in the singular frames (U, V). We verify this functionally in both directions: put the base spectrum back into an RL checkpoint at any mixing ratio — every benchmark stays at RL level. Graft the RL spectrum into pre-RL frames — nothing improves. </p>\n<p>ISO turns this into an optimization stack:<br>• <strong>ISO-Merger (offline)</strong>: compose shared-base RL experts with no data, rollouts, gradients, or distillation — beats TA/TIES/TSV/RAM/OrthoMerge on 7B and 1.5B suites.<br>• <strong>ISO-Optimizer (online)</strong>: run AdamW/Muon on (U, V) with Σ₀ frozen — matched RLVR accuracy in up to 2.7× fewer steps, on-par or higher final rewards.</p>\n<p>Interactive spectrum-swap demos (drag α yourself): <a href=\"https://iso-rlvr.github.io/\" rel=\"nofollow\">https://iso-rlvr.github.io/</a></p>\n","updatedAt":"2026-07-22T04:40:00.744Z","author":{"_id":"63884e0a3143f4706312f8eb","avatarUrl":"/avatars/cd9262a8b5b6e273c629b4d1c76a6d46.svg","fullname":"Wenyan Cong","name":"FOGmiaow","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.771722674369812},"editors":["FOGmiaow"],"editorAvatarUrls":["/avatars/cd9262a8b5b6e273c629b4d1c76a6d46.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.19331","authors":[{"_id":"6a60361c7e7f152167e47124","name":"Hanqing Zhu","hidden":false},{"_id":"6a60361c7e7f152167e47125","name":"Wenyan Cong","hidden":false},{"_id":"6a60361c7e7f152167e47126","name":"Zhizhou Sha","hidden":false},{"_id":"6a60361c7e7f152167e47127","name":"Sagnik Mukherjee","hidden":false},{"_id":"6a60361c7e7f152167e47128","name":"Xinyuan Song","hidden":false},{"_id":"6a60361c7e7f152167e47129","name":"David González-Martínez","hidden":false},{"_id":"6a60361c7e7f152167e4712a","name":"Xiaoxia Wu","hidden":false},{"_id":"6a60361c7e7f152167e4712b","name":"Yuandong Tian","hidden":false},{"_id":"6a60361c7e7f152167e4712c","name":"Shiwei Liu","hidden":false},{"_id":"6a60361c7e7f152167e4712d","name":"David Z. Pan","hidden":false},{"_id":"6a60361c7e7f152167e4712e","name":"Zhangyang \"Atlas\" Wang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/63884e0a3143f4706312f8eb/7Z2exJEaUm2MPrrsMsxNV.png","https://cdn-uploads.huggingface.co/production/uploads/63884e0a3143f4706312f8eb/FxiWNJTrApKUeUWzwhj7m.gif","https://cdn-uploads.huggingface.co/production/uploads/63884e0a3143f4706312f8eb/njcnmD8Nh-dUXmq8m2S3i.gif"],"publishedAt":"2026-07-21T17:51:36.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"ISO: An RLVR-Native Optimization Stack","submittedOnDailyBy":{"_id":"63884e0a3143f4706312f8eb","avatarUrl":"/avatars/cd9262a8b5b6e273c629b4d1c76a6d46.svg","isPro":false,"fullname":"Wenyan Cong","user":"FOGmiaow","type":"user","name":"FOGmiaow"},"summary":"Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames.\n We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.","upvotes":2,"discussionId":"6a60361c7e7f152167e4712f","projectPage":"https://iso-rlvr.github.io/","githubRepo":"https://github.com/zhuhanqing/ISO","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"620be1c49e55c0fe782f7f78","name":"UTEXAS","fullname":"University of Texas at Austin","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/OSAIQQGBT7YDemNgJlzHh.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63884e0a3143f4706312f8eb","avatarUrl":"/avatars/cd9262a8b5b6e273c629b4d1c76a6d46.svg","isPro":false,"fullname":"Wenyan Cong","user":"FOGmiaow","type":"user"},{"_id":"65905af887944e494e37e09a","avatarUrl":"/avatars/8f079ef18d06f506b94edca7d49e4c26.svg","isPro":false,"fullname":"Hert4","user":"beyoru","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"620be1c49e55c0fe782f7f78","name":"UTEXAS","fullname":"University of Texas at Austin","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/OSAIQQGBT7YDemNgJlzHh.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.19331.md","query":{}}">
Papers
arxiv:2607.19331

ISO: An RLVR-Native Optimization Stack

Published on Jul 21
· Submitted by
Wenyan Cong
on Jul 22
Authors:
,

Abstract

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.

Community

Paper submitter about 4 hours ago

🪐 ISO: Inherit the spectrum, optimize the frames.

The finding. RLVR barely rewrites the base model's singular-value spectra, and the new capability lives in the singular frames (U, V). We verify this functionally in both directions: put the base spectrum back into an RL checkpoint at any mixing ratio — every benchmark stays at RL level. Graft the RL spectrum into pre-RL frames — nothing improves.

ISO turns this into an optimization stack:
ISO-Merger (offline): compose shared-base RL experts with no data, rollouts, gradients, or distillation — beats TA/TIES/TSV/RAM/OrthoMerge on 7B and 1.5B suites.
ISO-Optimizer (online): run AdamW/Muon on (U, V) with Σ₀ frozen — matched RLVR accuracy in up to 2.7× fewer steps, on-par or higher final rewards.

Interactive spectrum-swap demos (drag α yourself): https://iso-rlvr.github.io/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.19331
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.19331 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.19331 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.19331 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers