Hugging Face Daily Papers · · 4 min read

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Introducing SAEScientist-Bench: Can AI agents conduct autonomous SAE interpretability research? We evaluate 10 agent configurations on 20 tasks covering feature discovery, activation selectivity, and steering. Agents find promising features, but steering remains challenging.<br>📄 <a href=\"https://arxiv.org/abs/2609.09113\" rel=\"nofollow\">https://arxiv.org/abs/2609.09113</a><br>💻 <a href=\"https://github.com/Trae1ounG/SAEScientist\" rel=\"nofollow\">https://github.com/Trae1ounG/SAEScientist</a></p>\n","updatedAt":"2026-09-10T04:28:37.021Z","author":{"_id":"64feb928c3329fa5933ebf9a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64feb928c3329fa5933ebf9a/DOyWiqah1dxjafj6TTGv5.jpeg","fullname":"TanYuQiao","name":"Trae1ounG","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.794751763343811},"editors":["Trae1ounG"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64feb928c3329fa5933ebf9a/DOyWiqah1dxjafj6TTGv5.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.09113","authors":[{"_id":"6aa0d2d0d0174964227bed04","name":"Yuqiao Tan","hidden":false},{"_id":"6aa0d2d0d0174964227bed05","name":"Shizhu He","hidden":false},{"_id":"6aa0d2d0d0174964227bed06","name":"Jun Zhao","hidden":false},{"_id":"6aa0d2d0d0174964227bed07","name":"Kang Liu","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?","submittedOnDailyBy":{"_id":"64feb928c3329fa5933ebf9a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64feb928c3329fa5933ebf9a/DOyWiqah1dxjafj6TTGv5.jpeg","isPro":false,"fullname":"TanYuQiao","user":"Trae1ounG","type":"user","name":"Trae1ounG"},"summary":"While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.","upvotes":14,"discussionId":"6aa0d2d0d0174964227bed08","githubRepo":"https://github.com/Trae1ounG/SAEScientist","githubRepoAddedBy":"user","ai_summary":"The study introduces a benchmark to evaluate AI agents using sparse autoencoders for autonomous mechanistic interpretability and feature discovery, revealing progress but significant gaps versus expert baselines.","ai_keywords":["Sparse Autoencoders","SAEs","mechanistic interpretability","contrastive probes","Gemma Scope","feature discovery","causal steering","activation rank","concept selectivity","autonomous AI R&D"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"640a887796aae649741a586f","name":"CASIA","fullname":"Chinese Academic of Science Institute of Automation","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1678411888885-6388984e8a5dbe2f3dc5afee.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64feb928c3329fa5933ebf9a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64feb928c3329fa5933ebf9a/DOyWiqah1dxjafj6TTGv5.jpeg","isPro":false,"fullname":"TanYuQiao","user":"Trae1ounG","type":"user"},{"_id":"68c4e7907daa73025f2b15ae","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/oabWp6ggSVwn5tLP4IMdz.jpeg","isPro":false,"fullname":"Boxuan Zhang","user":"ZBox008003","type":"user"},{"_id":"6a7161c357fc52b7454f9fe5","avatarUrl":"/avatars/93cc2d3e744785b09ad173304d37bb21.svg","isPro":false,"fullname":"x","user":"mixturex","type":"user"},{"_id":"694b72d0f71c3b988f1aacd4","avatarUrl":"/avatars/73c3d1a5107f74962da739a6b7b14263.svg","isPro":false,"fullname":"qiunan","user":"qiunannan","type":"user"},{"_id":"694b7486d7e02d8a1c1bcc2b","avatarUrl":"/avatars/edf0ce66c8bbafb8e13155de7d36a0f7.svg","isPro":false,"fullname":"mix leangth","user":"mixle","type":"user"},{"_id":"6a716221440f2693fdf2eb32","avatarUrl":"/avatars/d8d55487c35ac41522df24164eb7d7bc.svg","isPro":false,"fullname":"x","user":"aalphaM","type":"user"},{"_id":"653f1ef4aabbf15fc76a259c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653f1ef4aabbf15fc76a259c/1jJDeTOJaJIKQZ4g3i8V3.jpeg","isPro":false,"fullname":"LLLeo Li","user":"LLLeo612","type":"user"},{"_id":"668378432698e06471cfd4d8","avatarUrl":"/avatars/f6182eb9a5228891348a250bdf42c690.svg","isPro":false,"fullname":"shiyang li","user":"flashlizard","type":"user"},{"_id":"68cec03c6cf461f47ebb868b","avatarUrl":"/avatars/1951ae0b2be210005cad9c340fa7a1ae.svg","isPro":false,"fullname":"lucxhr","user":"lucxhr","type":"user"},{"_id":"6888610263bcb2b479ceff82","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/vi6ZbeD-gMU6iBiwwutaC.png","isPro":false,"fullname":"yangjingxiao","user":"yangjx29","type":"user"},{"_id":"6a30bc0f7246336616384101","avatarUrl":"/avatars/8a15598f1f35a2a16589e121902e0a5f.svg","isPro":false,"fullname":"wlings","user":"wlings","type":"user"},{"_id":"65a0aade5fafc248c2156e95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a0aade5fafc248c2156e95/S9YjJMTuKc-U1cFizqUMA.jpeg","isPro":false,"fullname":"DeyangKong","user":"DeyangKong","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"640a887796aae649741a586f","name":"CASIA","fullname":"Chinese Academic of Science Institute of Automation","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1678411888885-6388984e8a5dbe2f3dc5afee.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.09113.md","query":{}}">
Papers
arxiv:2609.09113

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Authors:
,

Abstract

The study introduces a benchmark to evaluate AI agents using sparse autoencoders for autonomous mechanistic interpretability and feature discovery, revealing progress but significant gaps versus expert baselines.

While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.

Community

Paper submitter about 4 hours ago

Introducing SAEScientist-Bench: Can AI agents conduct autonomous SAE interpretability research? We evaluate 10 agent configurations on 20 tasks covering feature discovery, activation selectivity, and steering. Agents find promising features, but steering remains challenging.
📄 https://arxiv.org/abs/2609.09113
💻 https://github.com/Trae1ounG/SAEScientist

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.09113
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.09113 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.09113 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.09113 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers