Research direction: Intelligent Model Weight transfer between LLMs [R]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Few days ago I feel like I need to get started with researching about LLMs. One thing which strikes the most in my mind , how we can reduce the time required for pre-training an LLM model to just few minutes. Right now the most efficient method that we have is knowledge distillation, which still takes time in response generation by the teacher model from prompts, backpropagation and training, to adjust the weights of student model making it to mimic the teacher model.
What if there is any way where we can adjust the model weights of an untrained model so that it becomes mathematically the same function as of the trained model.I want to figure out if there any such algorithm exist which would perform simple mathematical operations on the untrained model such that it becomes mathematically same function as the trained model.
If this become successful there is no need of training under distillation process or any conventional process, just few math operations on the untrained model, and then it's done, which would be taking few minutes. I need guidance and collaboration for someone who is working in this direction.
[link] [comments]
More from r/MachineLearning
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
-
worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P]
Aug 13
-
UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.