r/MachineLearning · · 1 min read

Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or secretly exhibit good behavior (which would be ironically here a misalignment, vs. misalignment behavior detected in current models) would that happen and would it be due to pre-training?

EDIT: What I am thinking if there is if there is some "alignment" already in pretraining (or like a raw latent machinery that alignment training later selects from) - would it show up in the naughty post-trained model eventually?

its late night so my brain is all over the place, but would love to hear your thoughts

submitted by /u/Objective_River_5218
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning