Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO.
There are a lot of papers on this, but because of limited compute I cannot try these papers out and learn them by implementing them myself.
If someone here has worked with these algorithms and their implementation on SLMs (something that can fit a consumer grade GPU like Nvidia RTX 4090 or 5090), can they suggest either a:
Github repo, or
The right choice of SLM(s) and the datasets, where i can see the difference between, RL/GRPO and OPSD algorithms?
Thanks in advance!
[link] [comments]
More from r/MachineLearning
-
A collision-entropy floor for watermark/retrieval AI-text detection. Looking for a sanity check before I take this further [D]
Aug 14
-
Are supervised and unsupervised learning still relevant today? [D]
Aug 14
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.