Hugging Face Daily Papers · · 6 min read

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We’re excited to share <strong><a href=\"https://arxiv.org/abs/2609.06289\" rel=\"nofollow\">Steering Geometry</a></strong>, accepted to <strong>EMNLP 2026 Main</strong>!</p>\n<p><strong>When we steer an LLM toward one human value, what happens to the others?</strong> We investigate whether steering directions capture the relationships predicted by psychological theories—and whether that structure translates into more consistent behavior.</p>\n<p>Our main contributions and findings:</p>\n<ul>\n<li>🧭 <strong>Geometry beyond steering accuracy:</strong> Across seven steering methods, distribution-driven approaches recover clearer human-value structure, while behavior-centric methods can achieve comparable steering performance with weak geometric alignment.</li>\n<li>🔄 <strong>Cross-value transfer:</strong> Better geometric alignment is associated with more theory-consistent transfer—steering toward one value strengthens compatible values and suppresses opposing ones.</li>\n<li>📈 <strong>Scale and instruction tuning:</strong> Value geometry improves with model scale but weakens after instruction tuning in our experiments.</li>\n<li>📊 <strong>Two value frameworks:</strong> We introduce a 26K-example contrastive benchmark covering 20 Schwartz values and a 1,200-example benchmark spanning six foundations of revised Moral Foundations Theory.</li>\n<li>🛠️ <strong>Tools for your own methods:</strong> Our code supports geometry analysis and cross-value transfer evaluation, making it straightforward to compare new steering interventions with the included methods.</li>\n</ul>\n<p>💻 <strong>Code:</strong> <a href=\"https://github.com/DeepRCL/Steering_Geometry\" rel=\"nofollow\">DeepRCL/Steering_Geometry</a><br>🤗 <strong>Dataset:</strong> <a href=\"https://huggingface.co/datasets/DeepRCL/SteeringGeometry\">DeepRCL/SteeringGeometry</a><br>📄 <strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2609.06289\" rel=\"nofollow\">Read on arXiv</a></p>\n<p>We’d love to hear your thoughts, especially on evaluating steering beyond the target behavior and extending this analysis to other concepts and value frameworks!</p>\n","updatedAt":"2026-09-09T04:05:13.747Z","author":{"_id":"64ba58d377dd483716aba098","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/6VASAUkFpDC-PR01yUJWj.png","fullname":"Mahdi Abootorabi","name":"aboots","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7608281373977661},"editors":["aboots"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/6VASAUkFpDC-PR01yUJWj.png"],"reactions":[],"isReport":false}},{"id":"6aa0e653a2082642fb0c0fb7","author":{"_id":"6aa0e5f3c250d3fd685b9e5b","avatarUrl":"/avatars/b081a9d10ca05ce121e3c9c9a020f09c.svg","fullname":"Arsolan abdb","name":"Arsalanabdol","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-09T04:53:39.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Very creative and good point of view","html":"<p>Very creative and good point of view</p>\n","updatedAt":"2026-09-09T04:53:39.101Z","author":{"_id":"6aa0e5f3c250d3fd685b9e5b","avatarUrl":"/avatars/b081a9d10ca05ce121e3c9c9a020f09c.svg","fullname":"Arsolan abdb","name":"Arsalanabdol","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9219793081283569},"editors":["Arsalanabdol"],"editorAvatarUrls":["/avatars/b081a9d10ca05ce121e3c9c9a020f09c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06289","authors":[{"_id":"6aa0d65bd0174964227bed55","user":{"_id":"64ba58d377dd483716aba098","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/6VASAUkFpDC-PR01yUJWj.png","isPro":false,"fullname":"Mahdi Abootorabi","user":"aboots","type":"user","name":"aboots"},"name":"Mohammad Mahdi Abootorabi","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.674Z","hidden":false},{"_id":"6aa0d65bd0174964227bed56","name":"Armin Saghafian","hidden":false},{"_id":"6aa0d65bd0174964227bed57","name":"Ali Bazshoushtari","hidden":false},{"_id":"6aa0d65bd0174964227bed58","name":"Hamid Rezaei","hidden":false},{"_id":"6aa0d65bd0174964227bed59","name":"EunJeong Hwang","hidden":false},{"_id":"6aa0d65bd0174964227bed5a","name":"Vered Shwartz","hidden":false},{"_id":"6aa0d65bd0174964227bed5b","name":"Parvin Mousavi","hidden":false},{"_id":"6aa0d65bd0174964227bed5c","name":"Purang Abolmaesumi","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64ba58d377dd483716aba098/43vHZLn0Y9Ak18ZIT-HXS.png","https://cdn-uploads.huggingface.co/production/uploads/64ba58d377dd483716aba098/2CgeFoYDfAAKM5YL0Kwtc.png"],"publishedAt":"2026-09-05T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"Steering Geometry: Validating Human Value Geometry in LLM Steering Space","submittedOnDailyBy":{"_id":"64ba58d377dd483716aba098","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/6VASAUkFpDC-PR01yUJWj.png","isPro":false,"fullname":"Mahdi Abootorabi","user":"aboots","type":"user","name":"aboots"},"summary":"As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman ρ up to 0.51, p < 10^{-13}). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.","upvotes":25,"discussionId":"6aa0d65cd0174964227bed5d","projectPage":"https://github.com/DeepRCL/Steering_Geometry","githubRepo":"https://github.com/DeepRCL/Steering_Geometry","githubRepoAddedBy":"user","ai_summary":"Activation steering vectors in large language models encode theory-aligned human value geometry when derived via distribution-driven methods, with geometric fidelity scaling with model size but declining after instruction tuning.","ai_keywords":["activation steering","steering vectors","latent geometry","Schwartz's Theory of Basic Human Values","distribution-driven methods","CAA","SphericalSteer","ODESteer","COLD-Steer","BiPO","instruction tuning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6a9b007e5ebd240ce7787841","name":"DeepRCL","fullname":"DeepRCL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/9Q7rXobS7NV6i2wLXaLi5.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64ba58d377dd483716aba098","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/6VASAUkFpDC-PR01yUJWj.png","isPro":false,"fullname":"Mahdi Abootorabi","user":"aboots","type":"user"},{"_id":"66f5bfeb1c540729cb300856","avatarUrl":"/avatars/96510703eb2fefffa55b5a9083e3bf54.svg","isPro":false,"fullname":"Nona Ghazizadeh","user":"nona-ghazizadeh","type":"user"},{"_id":"65dc61d9ad7ccf910d0c5df2","avatarUrl":"/avatars/ee691cd5a7dee29379d7a1e47569c452.svg","isPro":false,"fullname":"Armin Saghafian","user":"ArminSa","type":"user"},{"_id":"65197a6e2b4fffcb41fe6118","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/n_EL7IHyrrYm5Qjueki2a.jpeg","isPro":false,"fullname":"Iman Mohammadi","user":"ImanM02","type":"user"},{"_id":"6aa0d9d12fb42fe43203f776","avatarUrl":"/avatars/5c357c0c4a25caa69971d7b6eea2d95e.svg","isPro":false,"fullname":"Matin","user":"Matinnej","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6aa0e5f3c250d3fd685b9e5b","avatarUrl":"/avatars/b081a9d10ca05ce121e3c9c9a020f09c.svg","isPro":false,"fullname":"Arsolan abdb","user":"Arsalanabdol","type":"user"},{"_id":"67fd788d79c9bf6c8ef9c4d8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mq6YqAXV1Nfv4EWKi6JSW.png","isPro":false,"fullname":"Masih Beigi Rizi","user":"masihbr","type":"user"},{"_id":"68dd44f1a4671ccd661e668e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/IHEQi-oG-4MkMqNTWWSMy.png","isPro":false,"fullname":"Mohammad Abolnejadian","user":"theablemo","type":"user"},{"_id":"6a6aa32982de75b814ca70f8","avatarUrl":"/avatars/e0601087d18c99a0df158444d6734c77.svg","isPro":false,"fullname":"Richard Wilson","user":"Rapid-Richard","type":"user"},{"_id":"6a6c7f62ef16c968823e3b4d","avatarUrl":"/avatars/42a1f19bd0ad542978225ebdb71b944d.svg","isPro":false,"fullname":"Brian Wilson","user":"Brian-Wilson","type":"user"},{"_id":"6a6de8afcd50c6f8f59c8f11","avatarUrl":"/avatars/f8fd05ff9b2a6bc3f31955e17d292bdb.svg","isPro":false,"fullname":"Sarah Miller","user":"Nimbus-Sarah","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a9b007e5ebd240ce7787841","name":"DeepRCL","fullname":"DeepRCL","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ba58d377dd483716aba098/9Q7rXobS7NV6i2wLXaLi5.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06289.md","query":{}}">
Papers
arxiv:2609.06289

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

Published on Sep 5
· Submitted by
Mahdi Abootorabi
on Sep 9

Abstract

Activation steering vectors in large language models encode theory-aligned human value geometry when derived via distribution-driven methods, with geometric fidelity scaling with model size but declining after instruction tuning.

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman ρ up to 0.51, p < 10^{-13}). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.

Community

Paper author Paper submitter about 10 hours ago edited about 10 hours ago

We’re excited to share Steering Geometry, accepted to EMNLP 2026 Main!

When we steer an LLM toward one human value, what happens to the others? We investigate whether steering directions capture the relationships predicted by psychological theories—and whether that structure translates into more consistent behavior.

Our main contributions and findings:

  • 🧭 Geometry beyond steering accuracy: Across seven steering methods, distribution-driven approaches recover clearer human-value structure, while behavior-centric methods can achieve comparable steering performance with weak geometric alignment.
  • 🔄 Cross-value transfer: Better geometric alignment is associated with more theory-consistent transfer—steering toward one value strengthens compatible values and suppresses opposing ones.
  • 📈 Scale and instruction tuning: Value geometry improves with model scale but weakens after instruction tuning in our experiments.
  • 📊 Two value frameworks: We introduce a 26K-example contrastive benchmark covering 20 Schwartz values and a 1,200-example benchmark spanning six foundations of revised Moral Foundations Theory.
  • 🛠️ Tools for your own methods: Our code supports geometry analysis and cross-value transfer evaluation, making it straightforward to compare new steering interventions with the included methods.

💻 Code: DeepRCL/Steering_Geometry
🤗 Dataset: DeepRCL/SteeringGeometry
📄 Paper: Read on arXiv

We’d love to hear your thoughts, especially on evaluating steering beyond the target behavior and extending this analysis to other concepts and value frameworks!

Very creative and good point of view

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06289
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.06289 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.06289 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers