A Method for Learning Value Systems in Generative AI
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computers and Society
Title:A Method for Learning Value Systems in Generative AI
Abstract:Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observing human behaviour. This work addresses the lack of grounded value learning methods in generative AI: existing approaches typically replicate human preferences without awareness of the multidimensional structure of value alignment, or lack principled value system elicitation methods. To address these gaps, we adapt a previously validated value system learning method to the generative AI setting, which, based on pairwise prompt-response preference data, simultaneously learns: i) an implementation of a grounding for a set of values given by a multi-objective reward model, and ii) a value system representation in the form of a weighted linear scalarization of the previous grounding model. To ensure that the learned value systems are based on coherent value representations, our algorithm dynamically prioritizes the grounding learning process. We evaluate the method against baselines and a contemporary method on prompt-response preference datasets. Results show competitive performance and minimal trade-offs against the baselines, while improving explainability.
| Comments: | Full version of a to be published paper in proceedings of the 9th AAAI/ACM conference in AI, Ethics and Society (AIES 2026). Includes supplementary material. 20 pages, 2 figures |
| Subjects: | Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.2.6; J.4; K.4.0 |
| Cite as: | arXiv:2607.16903 [cs.CY] |
| (or arXiv:2607.16903v1 [cs.CY] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16903
arXiv-issued DOI via DataCite (pending registration)
|
Submission history
From: Andrés Holgado-Sánchez [view email][v1] Sat, 18 Jul 2026 17:37:50 UTC (393 KB)
Access Paper:
- View PDF
- TeX Source
Current browse context:
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
Aug 7
-
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
Aug 7
-
Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs
Aug 7
-
GROM: Gradient-Free Rapid One-Shot Machine Unlearning
Aug 7
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.