Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Setup for one long agentic coding session (tools + vision) on a single 5090: Qwen3.8-27B RVN Heretic (ARA abliterated) at Q6_K, 159744 context, 1 slot, q8_0 K/V, DFlash2 speculative decoding, vision projector on the GPU, Sharp chat template. All numbers measured 2026-09-16.
Weights: 0bserverx RVN, built on trohrbaugh/heretic-ara (KL 0.0535, 3/100 refusals) plus two more ARA passes: KL 0.0085 vs base, 0-1/100 refusals. The Q6_K file has no MTP head; DFlash2 is the spec path.
Results (greedy, thinking off, 300-token generations)
- Decode with DFlash2 n=4: 140-190 tok/s on code depending on the prompt, ~100 on prose (acceptance decides it). Same model with no speculation: 62 tok/s.
- Cold prefill (measured without the drafter): 32K in 11.7 s, 64K in 28.6 s, ~113K in 60 s.
- VRAM: 29996 MiB after load, 30332 MiB after the first image, with a desktop holding ~1 GB. 163840 loads but leaves too little for the first image.
- Tool calls parse, images read correctly, no Xid.
Downloads
| What | File | Link |
|---|---|---|
| Target + vision | RVN-Q6_K-multilingual.gguf (22.08 GB), mmproj-Qwen3.8-27B-Q8_0.gguf (600 MiB) | repo |
| Draft | Qwen3.8-27B-DFlash2-Q4_K_M.gguf (1.14 GB) | z-lab/Qwen3.8-27B-DFlash2-GGUF |
| Chat template | chat_template.jinja, Sharp v22.5.0 (root file is whatever is latest, check the README) | peculiar-ragdoll/Qwen-Sharp-Chat-Templates |
Target sha256: 591be92b29db500eca2c75f8d0f292260426814418e50a8bf2f6e8a1ed29d138 |
Engine: llama.cpp master as of 2026-09-16. Needs DFlash2 (#27342) and the image speculation fix (#28715), both merged. My build also carried the open chunked GDN prefill PR (#26001); it only affects prefill, decode numbers reproduce on plain master.
hf download 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF \ RVN-Q6_K-multilingual.gguf mmproj-Qwen3.8-27B-Q8_0.gguf --local-dir ./rvn hf download z-lab/Qwen3.8-27B-DFlash2-GGUF Qwen3.8-27B-DFlash2-Q4_K_M.gguf --local-dir ./dflash2 hf download peculiar-ragdoll/Qwen-Sharp-Chat-Templates chat_template.jinja --local-dir ./sharp Serve
GGML_CUDA_DISABLE_GRAPHS=1 llama-server \ --model ./rvn/RVN-Q6_K-multilingual.gguf \ --mmproj ./rvn/mmproj-Qwen3.8-27B-Q8_0.gguf --mmproj-offload \ --model-draft ./dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf \ --spec-type draft-dflash,ngram-map-k4v --spec-draft-n-max 4 \ --spec-ngram-map-k4v-size-n 12 --spec-ngram-map-k4v-size-m 48 --spec-ngram-map-k4v-min-hits 1 \ --spec-draft-ngl all --n-gpu-layers all --fit off \ --ctx-size 159744 --batch-size 2048 --ubatch-size 512 --parallel 1 \ --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \ --cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \ --cache-ram 16384 \ --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 \ --image-min-tokens 1024 --image-max-tokens 2048 \ --jinja --chat-template-file ./sharp/chat_template.jinja \ --reasoning-format auto --reasoning-preserve \ --host 127.0.0.1 --port 8888 --alias qwen3.8-27b Notes:
- CUDA graphs are off because of #27330 (Xid 8 hangs on the 5090).
--cache-ram 16384keeps up to 16 GiB of parked prompt states in system RAM so a long conversation is not re-prefilled.- Do not add
--swa-full: the DFlash2 drafter's 5 layers are sliding-window 2048, so its KV stays ~40 MiB at any context. ngram-map-k4vis stacked on DFlash2 for free (copy-from-context in agentic coding). If prose gets slower, use--spec-type draft-dflashalone.- Token embeddings stay on the host, which is where the last ~1 GB of context comes from.
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.