r/LocalLLaMA · · 1 min read

Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro

Settings are as follows

--n-gpu-layers 999 \

--n-cpu-moe 37 \

--no-mmap \

-ctk q8_0 \

-ctv q8_0 \

-fa 1 \

-c 9000 \

submitted by /u/Sweaty_Perception655
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA