Hugging Face Daily Papers · · 5 min read

Scaling Inherently Interpretable Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<strong>Technical Report</strong></p>\n<p>Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale.<br>We instantiate the training-time recipe with <strong>Steerling-8B</strong>, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. <strong>Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.</strong></p>\n","updatedAt":"2026-08-11T05:52:48.334Z","author":{"_id":"61fb0c7bb3d6dbddda6dbe44","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1643842677900-noauth.jpeg","fullname":"Anonymous","name":"luulinh90s","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9141676425933838},"editors":["luulinh90s"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1643842677900-noauth.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07594","authors":[{"_id":"6a7ab82b019ce76dc7b3ab31","name":"Guide Labs Team","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab32","name":"Andreas Madsen","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab33","name":"Aya Abdelsalam Ismail","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab34","name":"Giang Nguyen","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab35","name":"Isaac Plant","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab36","name":"Muawiz Chaudhary","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab37","name":"Nathaniel Monson","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab38","name":"Saqib Azim","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab39","name":"Zhichen Guo","hidden":false},{"_id":"6a7ab82b019ce76dc7b3ab3a","name":"Julius Adebayo","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Scaling Inherently Interpretable Language Models","submittedOnDailyBy":{"_id":"61fb0c7bb3d6dbddda6dbe44","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1643842677900-noauth.jpeg","isPro":false,"fullname":"Anonymous","user":"luulinh90s","type":"user","name":"luulinh90s"},"summary":"Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale.\n We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.","upvotes":13,"discussionId":"6a7ab82c019ce76dc7b3ab3b","projectPage":"https://www.guidelabs.ai/papers/scaling-inherently-interpretable-language-models/","githubRepo":"https://github.com/guidelabs/steerling","githubRepoAddedBy":"user","ai_summary":"Integrating interpretability as a training constraint yields scalable, disentangled representations that enable attribution, retrieval, and steering without retraining.","ai_keywords":["autoregressive language models","diffusion language models","causal attention mask","disentangled representations","concept attribution","feature attribution","training data retrieval","concept steering","scaling paradigm"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":238,"organization":{"_id":"667c4fe11478793f9e89c2ca","name":"guidelabs","fullname":"Guide Labs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/663e6f91b2f118e6199bd1e4/AtwRZCYIlbAjiEnGLl9VV.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"61fb0c7bb3d6dbddda6dbe44","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1643842677900-noauth.jpeg","isPro":false,"fullname":"Anonymous","user":"luulinh90s","type":"user"},{"_id":"6618413e625ae15f64769476","avatarUrl":"/avatars/1585703eab341f7808e7bf703ffef9dc.svg","isPro":false,"fullname":"Minh-Quan Le","user":"lmquan","type":"user"},{"_id":"63162ef093ab42acfb060200","avatarUrl":"/avatars/9d5bc4e5070b3e1d5a18da2ebbe90a7a.svg","isPro":false,"fullname":"Julius Adebayo","user":"juliusad","type":"user"},{"_id":"68f6d167584b6ff1c6ba361d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f6d167584b6ff1c6ba361d/SJlQTnvWkMFns7qfF4PK9.jpeg","isPro":false,"fullname":"Giang Nguyen","user":"giangnguyen7-glai","type":"user"},{"_id":"6723f0242fd598f719901405","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/dEBCSQG48oyErejoTsqQr.jpeg","isPro":false,"fullname":"Aya Abdelsalam Ismail","user":"asalam91","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"60b63a757430e735fbe737b1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1622555235160-noauth.jpeg","isPro":false,"fullname":"Andreas Madsen","user":"andreasmadsen","type":"user"},{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"},{"_id":"653a169e7174042dab9ab090","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/RcPzQ-CbPCldhJcitHJ6N.jpeg","isPro":false,"fullname":"Siyuan He","user":"Siyuan0730","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"64938a635b1cbc4829d7fa92","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64938a635b1cbc4829d7fa92/rQg55CnsJfB3y1uVBXnAr.jpeg","isPro":false,"fullname":"Saqib Azim","user":"saqib1707","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"667c4fe11478793f9e89c2ca","name":"guidelabs","fullname":"Guide Labs","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/663e6f91b2f118e6199bd1e4/AtwRZCYIlbAjiEnGLl9VV.jpeg"},"query":{}}">
Papers
arxiv:2608.07594

Scaling Inherently Interpretable Language Models

Published on Aug 6
· Submitted by
Anonymous
on Aug 11
Authors:
,

Abstract

Integrating interpretability as a training constraint yields scalable, disentangled representations that enable attribution, retrieval, and steering without retraining.

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale. We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.

Community

Paper submitter about 13 hours ago

Technical Report

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale.
We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.07594 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.07594 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.07594 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers