I made a tool that chains a small local model into a big coding model and auto-unloads VRAM between them
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| A couple weeks ago I shared PromptChain here a small Streamlit app that chains two models: a little Prompter that rewrites your rough idea into a proper prompt, then a larger Coder that turns that prompt into code. The whole point is that on an 8–16 GB card you can usually only hold one model at a time, so it auto-unloads one before loading the other no manual swapping, no copy-pasting between two chat windows. The comments last time turned into a real to-do list, so here's what's landed since:
The part I still like most: keep the Prompter local and point the Coder at a cloud model (OpenAI/Claude/Gemini). You fix the prompt for free on the local model, so the one paid generation lands right more often and you re-roll way less — frontier code quality without paying for every re-roll. Local-first, MIT, no telemetry. Works with LM Studio, Ollama, or any OpenAI-compatible server GitHub: Genuinely after feedback both positive and negative. Also feel free to tell Prompter/Coder pairings that work well on your hardware [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.