Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.
Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.
The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under ~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys ~1GB of VRAM.
Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.
Launch command:
bat
@echo off
cd /d "%~dp0"
"%~dp0llama-server.exe" ^
-m "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf" ^
--mmproj "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" ^
--no-mmproj-offload ^
-ngl 99 ^
--n-cpu-moe 39 ^
-c 131072 ^
-np 1 ^
-t 6 ^
-tb 10 ^
-b 2048 ^
-ub 2048 ^
-fa on ^
-ctk q8_0 ^
-ctv q8_0 ^
--load-mode mmap+mlock ^
--jinja ^
--reasoning-format deepseek ^
--reasoning-preserve ^
--spec-type none ^
--image-min-tokens 1024 ^
--temp 0.6 ^
--top-p 0.95 ^
--top-k 20 ^
--min-p 0 ^
--alias Qwen3.6-35B-A3B-Uncensored-Q4_K_M ^
--host 127.0.0.1 ^
--port 8081
pause
Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU
Context: 131072 tokens, Q8_0 KV cache
Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.