actual advice about SLM fine tuning?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
hello real people and less-real bots, i'd appreciate if any of you people who have fine-tuned (either full or peft) more than half a model could share your wisdom about fine-tuning. i know i can ask the friendly neighborhood chatgpt and also unsloth has some detailed docs but that's not what i'm looking for. i'm looking for things like
- here's a good way to think about curating a dataset, or
- that-and-that lora rank suit that-and-that task, or
- if your cost is looking like <that> maybe check the gradients, it happened to me once because because of <that>, or
- if you want to train for <something> you should start by training for <simple thing> and gradually move towards <harder thing>
- you should really think about <that> when designing your benchmarks
i think most of you are smart enough to understand where this is heading. i'm interested in both introducing new knowledge and improving abilities on a less-common language.
EDIT: following u/Environmental-Metal9's comment (thanks!), i wanted to clarify: let's say i have this SLM that knows a bit about marine biology (for example) and i want it to become very knowledgeable about it. why? because i want to integrate it into this pipeline i build for some company in that domain, because i want to use it as a llm-as-a-judge, or just because i want to learn fine-tuning. i still want it to maintain its reasoning capabilities, though. so i know i'll probably need to do through some SFT and alignment stages, fine (best practices for generating synth data?), i don't mind starting small with a "bit better" model that knows a little more about fish but doesn't know what's the color of the sky.
any advice is welcome and apologies for my lazy english. not letting gpt to correct me though, it's nice to see some less polished texts every now and then i think (and people also like to comment about typos, so i have learned).
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.