r/LocalLLaMA · · 1 min read

I built a cache-friendly context compacting plugin for OpenCode

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

https://github.com/lennartschoch/opencode-cache-compact

The default context compacting mechanism in OpenCode strips a bunch of tokens from the beginning of the conversation (system prompt, tools etc).

This is fine for hosted models, but on a local model this means you'll prefill the entire conversation that's already cached.

I built a plugin that keeps the conversation as-is, prompts the model to write a summary and then transforms the conversation to erase everything aside from system prompt, tools and summary - because the entire conversation is cached, this is super fast (usually around 1-2 minutes on my Strix Halo, previously >10min).

Would love to get some feedback on this - is this useful for anyone else? This is my first open source project in the local LLM space so I'd love to know your thoughts!

submitted by /u/schennardo
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA