r/LocalLLaMA · · 2 min read

Qwen3.8-27B uncensored Q6_K at 156K context on one RTX 5090, 140-190 tok/s with DFlash2

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Setup for one long agentic coding session (tools + vision) on a single 5090: Qwen3.8-27B RVN Heretic (ARA abliterated) at Q6_K, 159744 context, 1 slot, q8_0 K/V, DFlash2 speculative decoding, vision projector on the GPU, Sharp chat template. All numbers measured 2026-09-16.

Weights: 0bserverx RVN, built on trohrbaugh/heretic-ara (KL 0.0535, 3/100 refusals) plus two more ARA passes: KL 0.0085 vs base, 0-1/100 refusals. The Q6_K file has no MTP head; DFlash2 is the spec path.

Results (greedy, thinking off, 300-token generations)

  • Decode with DFlash2 n=4: 140-190 tok/s on code depending on the prompt, ~100 on prose (acceptance decides it). Same model with no speculation: 62 tok/s.
  • Cold prefill (measured without the drafter): 32K in 11.7 s, 64K in 28.6 s, ~113K in 60 s.
  • VRAM: 29996 MiB after load, 30332 MiB after the first image, with a desktop holding ~1 GB. 163840 loads but leaves too little for the first image.
  • Tool calls parse, images read correctly, no Xid.

Downloads

What File Link
Target + vision RVN-Q6_K-multilingual.gguf (22.08 GB), mmproj-Qwen3.8-27B-Q8_0.gguf (600 MiB) repo
Draft Qwen3.8-27B-DFlash2-Q4_K_M.gguf (1.14 GB) z-lab/Qwen3.8-27B-DFlash2-GGUF
Chat template chat_template.jinja, Sharp v22.5.0 (root file is whatever is latest, check the README) peculiar-ragdoll/Qwen-Sharp-Chat-Templates
Target sha256: 591be92b29db500eca2c75f8d0f292260426814418e50a8bf2f6e8a1ed29d138

Engine: llama.cpp master as of 2026-09-16. Needs DFlash2 (#27342) and the image speculation fix (#28715), both merged. My build also carried the open chunked GDN prefill PR (#26001); it only affects prefill, decode numbers reproduce on plain master.

hf download 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF \ RVN-Q6_K-multilingual.gguf mmproj-Qwen3.8-27B-Q8_0.gguf --local-dir ./rvn hf download z-lab/Qwen3.8-27B-DFlash2-GGUF Qwen3.8-27B-DFlash2-Q4_K_M.gguf --local-dir ./dflash2 hf download peculiar-ragdoll/Qwen-Sharp-Chat-Templates chat_template.jinja --local-dir ./sharp 

Serve

GGML_CUDA_DISABLE_GRAPHS=1 llama-server \ --model ./rvn/RVN-Q6_K-multilingual.gguf \ --mmproj ./rvn/mmproj-Qwen3.8-27B-Q8_0.gguf --mmproj-offload \ --model-draft ./dflash2/Qwen3.8-27B-DFlash2-Q4_K_M.gguf \ --spec-type draft-dflash,ngram-map-k4v --spec-draft-n-max 4 \ --spec-ngram-map-k4v-size-n 12 --spec-ngram-map-k4v-size-m 48 --spec-ngram-map-k4v-min-hits 1 \ --spec-draft-ngl all --n-gpu-layers all --fit off \ --ctx-size 159744 --batch-size 2048 --ubatch-size 512 --parallel 1 \ --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \ --cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \ --cache-ram 16384 \ --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 \ --image-min-tokens 1024 --image-max-tokens 2048 \ --jinja --chat-template-file ./sharp/chat_template.jinja \ --reasoning-format auto --reasoning-preserve \ --host 127.0.0.1 --port 8888 --alias qwen3.8-27b 

Notes:

  • CUDA graphs are off because of #27330 (Xid 8 hangs on the 5090).
  • --cache-ram 16384 keeps up to 16 GiB of parked prompt states in system RAM so a long conversation is not re-prefilled.
  • Do not add --swa-full: the DFlash2 drafter's 5 layers are sliding-window 2048, so its KV stays ~40 MiB at any context.
  • ngram-map-k4v is stacked on DFlash2 for free (copy-from-context in agentic coding). If prose gets slower, use --spec-type draft-dflash alone.
  • Token embeddings stay on the host, which is where the last ~1 GB of context comes from.
submitted by /u/Fz1zz
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA