r/LocalLLaMA · · 1 min read

Combining RAG with Continued Pretraining

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

For leaning purposes, I ran an experiment where I trained a model on a new domain using continued pretraining. Then to make it more flexible, I added a RAG step to inject dynamic data to augment the stable training data.

The specific example is training qwen 3.5 4B on a fictional subway system to where the model learns the map well enough to provide travel routes (including multiple transfers, etc). A key point here is to avoid a training corpus that relies on memorization.

Once the subway map was stable and generalized through CPT, I added a RAG step to support dynamic travel announcements (e.g. station closure, concerts near a station, etc).

It's definitely been a fun experiment. Check it out here in case you are interested:

https://www.teachmecoolstuff.com/viewarticle/combining-rag-with-continued-pretraining-of-llms

submitted by /u/funJS
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA