r/LocalLLaMA · · 1 min read

How are you guys thinking about context now, and building around it?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Not asking for anyone’s secrets of the trade, I’m more curious how people are thinking about context now that newer models chew through huge amounts of it for reasoning.

the TLDR: I’m starting to think of context less as working memory and more as a temp scratchpad to start each step.

I’m running a small setup: 32gb vram on my main PC, and an older machine with 8GB running a 9B Qwen model in the background as a compaction and long-term-memory sorter.

My main model’s working state lives outside the context window in docs that it continuously writes and edits. The context has become more about whatever it needs for the current task, plus retrieval from those docs when needed with git there for recall and history.

So I'm just trying to gauge where other people on the lower end of local hosting have landed with this. especially without throwing in bloated systems for supporting it.

submitted by /u/Training-Ruin-5287
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA