LARA: small, composable behaviours for frozen LLMs [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| GitHub: https://github.com/pfekin/LARA I've been working on LARA (Lightweight Additive Residual Adaptation), a research project on making post-training modular for frozen language models. I've also developed a small PyTorch library that implements it. The main idea is to train a low-rank residual adapter at selected layers rather than modifying the model's weights. The resulting behaviors are small enough to keep separately and can be loaded, removed, blended or routed at inference time. The Mixture of Behaviors (MoBs) demo is maybe the easiest way to see what this means in practice. Several independently trained behaviors can share the same frozen model, with a soft router selecting or combining them on a token by token basis. For example, a single model can have separate coding, maths, medical and summmarization behaviors rather than keeping four separately adapted models. The repository also includes a comparison with LoRA and some writing style behaviors trained on Hemingway, Fitzgerald and Gertrude Stein (second demo). It's an on-going research project, but the library is usable now and includes the training code, examples and reproduction instructions (for the paper). [link] [comments] |
More from r/MachineLearning
-
Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]
Sep 28
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.