Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or secretly exhibit good behavior (which would be ironically here a misalignment, vs. misalignment behavior detected in current models) would that happen and would it be due to pre-training?
EDIT: What I am thinking if there is if there is some "alignment" already in pretraining (or like a raw latent machinery that alignment training later selects from) - would it show up in the naughty post-trained model eventually?
its late night so my brain is all over the place, but would love to hear your thoughts
[link] [comments]
More from r/MachineLearning
-
For the people who got reviews back from neurips, cvpr, eccv, etc and also tested their paper through an agentic reviewer like the stanford one, how different were the reviews? [D]
Aug 14
-
Building text to ASCII diffusion model , need advice and guidance [P]
Aug 14
-
A collision-entropy floor for watermark/retrieval AI-text detection. Looking for a sanity check before I take this further [D]
Aug 14
-
Are supervised and unsupervised learning still relevant today? [D]
Aug 14
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.