r/LocalLLaMA · · 1 min read

Is llama.cpp meant to be slow at long context, even when you aren't using that context?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I am trying out a few fine-tunes of Qwen 3.5 9B @ IQ4_XS @ 131K context and trying to go mostly local (free beats cheap, after all). However, it is much slower than at, say 16K context, even when I am not actually using 131K tokens in the first place. Anyone know why this is? I am using the following command on an 8 GB laptop 4060 (I've tried to set aside ~1GB for the OS and whatnot).

llama serve -hf bartowski/Ornith-1.5-9B-GGUF:IQ4_XS --fit on --cache-type-k q8_0 --cache-type-v q4_1 -c 131072 --temp 0.7

submitted by /u/Aggravating-Push-207
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA