r/LocalLLaMA · · 2 min read

Running qwen 3.8 27B iq3 xxs on RTX 3060.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Running qwen 3.8 27B iq3 xxs on RTX 3060.

Getting anywhere from 10 - 20 tps.
Thinking Off . Took about 4 mins and 7 mins.
Running on about "IQ3_S - 3.4375 bpw"

37.03.960.932 I slot print_timing: id 0 | task 2665 | prompt processing, n_tokens = 12516, progress = 0.98, t = 37.08 s / 337.55 tokens per second
37.04.777.411 I slot print_timing: id 0 | task 2665 | prompt processing, n_tokens = 12768, progress = 1.00, t = 37.78 s / 337.94 tokens per second
37.13.746.606 I slot print_timing: id 0 | task 2665 | n_gen = 100, tg = 11.34 t/s, tg_3s = 11.45 t/s
37.16.955.263 I slot print_timing: id 0 | task 2665 | n_gen = 140, tg = 11.64 t/s, tg_3s = 12.47 t/s
37.20.166.770 I slot print_timing: id 0 | task 2665 | n_gen = 180, tg = 11.81 t/s, tg_3s = 12.46 t/s
37.23.366.284 I slot print_timing: id 0 | task 2665 | prompt eval time = 38241.22 ms / 12772 tokens ( 2.99 ms per token, 333.99 tokens per second)
37.23.366.290 I slot print_timing: id 0 | task 2665 | eval time = 18350.47 ms / 211 tokens ( 87.38 ms per token, 11.44 tokens per second)
37.23.366.291 I slot print_timing: id 0 | task 2665 | total time = 56591.69 ms / 12983 tokens
37.23.366.292 I slot print_timing: id 0 | task 2665 | graphs reused = 2342
37.23.366.296 I slot print_timing: id 0 | task 2665 | draft acceptance = 0.41250 ( 132 accepted / 320 generated), mean len = 2.65
37.23.366.758 I slot release: id 0 | task 2665 | stop processing: n_tokens = 12984, truncated = 0

My run for this is

~/sandbox/dcfr/third_party/llama.cpp/build-cuda/bin main* ❯ export LLAMA_GDN_TRANSACTIONAL_REPLAY=1 export GGML_OP_OFFLOAD_MIN_BATCH=2 ./llama-server \ -m /mnt/D/Mymodels/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \ --alias qwen3.8-27b-iq3-64k-dcfr \ -c 65536 \ --parallel 1 \ -dev CUDA0 \ --fit off \ --n-gpu-layers 50 \ --override-tensor 'blk\.(10|11|12|13|14|15|16)\..*=CUDA0' \ --load-mode none \ -ctk q4_0 \ -ctv q4_0 \ -b 256 \ -ub 256 \ -t 6 \ -tb 6 \ --spec-type draft-mtp \ --spec-draft-n-max 4 \ --spec-draft-p-min 0 \ --spec-draft-type-k q4_0 \ --spec-draft-type-v q4_0 \ --spec-draft-threads 6 \ --spec-draft-threads-batch 6 \ -fa on \ --no-mmproj \ --reasoning off \ --jinja \ --cache-ram 128 \ --no-cache-idle-slots \ --no-ui \ --host 127.0.0.1 \ --port 5800 \ --metrics 

I still have more than 1GB vram left after loading the full model and kv cache.
Or you can follow his guide https://github.com/kadenball/qwen38-27b-rtx3060-dcfr

If anybody got issues running on rtx 3060 tell me.

submitted by /u/SummarizedAnu
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA