r/LocalLLaMA · · 6 min read

I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8...

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practical winner so far is a recent CUDA build of llama.cpp with tensor split, Flash Attention, --numa distribute, and large batches. On Qwen3.8 27B Q6_K_M I measured 1,376.9 prompt tok/s at 2k, 1,324.3 prompt tok/s at 4k, 1,221.5 prompt tok/s at 16k, and 39.9 decode tok/s. I also made Qwen3.8 Flash Next Q4 GGUF of a 177 B/6B-active MoE model run across the same two cards at roughly 26 tok/s decode and 380 PP.

This is not a claim that old V100s beat modern GPUs. It is a report of what they can do in this crazy "a kidney for a GPU" market!

My Hardware:

| Layer | Hardware / configuration | |---|---| | GPU 0 | NVIDIA Tesla V100-PCIE, 16,144 MiB VRAM, compute capability 7.0 | | GPU1 | NVIDIA Tesla V100-PCIE, 32,494 MiB VRAM, compute capability 7.0 | | Server topology | NUMA-aware host; benchmarks use `numactl --interleave=all` or `--numa distribute` | | Virtualization | Proxmox with LXC inference containers for isolated runtimes | | CPU| 2 x Xeon E5-2696 v4 @ 2.20GHz| | RAM | 512 GB DDR4 2400 ECC RDIMM| | CUDA stack | CUDA 12.8; NVIDIA driver libraries are host-mounted into the vLLM LXC where needed | 
Workload Runtime / model Best result Configuration behind the result
Best 27B prefill Mainline llama.cpp, Qwen3.8 27B Q6_K_M 1,376.87 pp tok/s at 2k Tensor split; main GPU = 32 GB V100; Flash Attention; Q8 KV; batch/ubatch 2,048; 20 threads; NUMA distribute
Best 27B decode Mainline llama.cpp, Qwen3.8 27B Q6_K_M 39.88 tg tok/s Same configuration
27B at 16k context Mainline llama.cpp, Qwen3.8 27B Q6_K_M 1,221.49 pp tok/s Same configuration
Best Q8 27B prefill Mainline llama.cpp, Qwen3.8 27B Q8_K_XL 1,059.67 pp tok/s at 2k Tensor split; long-context run used 1:1.5 split
Q8 27B at 64k Mainline llama.cpp, Qwen3.8 27B Q8_K_XL 664.36 pp tok/s 65,536-token prompt; tensor split 1:1.5
Best ik_llama 27B result ik_llama.cpp, Qwen3.8 27B Q6_K_M 670.61 pp tok/s at 2k; 35.43 tg tok/s Graph split; 1:1 split; smaller batch settings than the mainline winner
Largest model tested ik_llama.cpp, Qwen3.8 Flash Next Q4_K_XL 380.67 pp tok/s at 4k; 26.29 tg tok/s 103.68 GiB GGUF; reported 177B total / 6B active MoE; graph split 1:1.75
ninfer-v100 NInfer, Qwen3.8 27B NVFP4 927.54 pp tok/s at 8,190 19.7 GiB weights; INT8 KV; MTP

What the tuning showed

Qwen3.8 27B Q6_K_M, mainline llama.cpp pp2048 tg128
Layer split, 79 threads 981.79 22.88
Tensor split, 20 threads 1,371.57 38.79
Tensor split, main GPU = 32 GB V100 1,376.87 39.88
Tensor split 1:2 980.71 26.10
No split 967.36 18.05

Takeaway: on this uneven V100 pair, tensor split—not layer split—and using the 32 GB V100 as main GPU are decisive. The best daily-driver path is mainline llama.cpp + Qwen3.8 27B Q6 + tensor split + Flash Attention + Q8 KV + NUMA distribution.

root@llama-cpp:~# numactl --interleave=all /opt/llama.cpp/build/bin/llama-bench -m /path/to/Qwen3.8-27B-UD-Q6_K_M.gguf -ngl 999 -sm tensor -mg 1 -fa on -b 2048 -ub 2048 -t 20 -p 512,2048,4096,16384 -n 128 -r

5 -lm none --numa distribute

ggml_cuda_init: found 2 CUDA devices (Total VRAM: 48638 MiB):

Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB

Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB

| model | size | params | backend | ngl | threads | n_ubatch | main_gpu | sm | fa | lm | test | t/s |

| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | -------: | ---------: | -----: | --: | ---------: | --------------: | -------------------: |

| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp512 | 1156.94 ± 5.10 |

| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp2048 | 1376.87 ± 4.59 |

| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp4096 | 1324.34 ± 7.37 |

| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp16384 | 1221.49 ± 3.87 |

| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | tg128 | 39.88 ± 0.04 |

build: b29c606 (1)

root@llama-cpp:~# for MAIN in 0 1; do

echo

echo "===== MAIN GPU ${MAIN} ====="

/usr/bin/numactl --interleave=all \

/opt/ik_llama.cpp/build/bin/llama-bench \

-m /path/to/models/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 -sm graph -ts 1/1.75 -mg "$MAIN" --fit 1 --fit-margin 1024 -fa 1 -p 4096,8192 -n 128 -b 2048 -ub 1024 -r 10 -o md
done 2>&1 | tee /root/bench-qwen38-main-gpu-final.md
===== MAIN GPU 0 =====

ggml_cuda_init: found 2 CUDA devices:
Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB
Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB

| model | size | params | backend | ngl | n_ubatch | sm | ts | test | t/s |

| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | ----: | ------------ | ------------: | ---------------: |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | pp4096 | 356.20 ± 7.68 |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | pp8192 | 361.00 ± 4.72 |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | tg128 | 26.44 ± 0.11

===== MAIN GPU 1 =====

ggml_cuda_init: found 2 CUDA devices:

Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB

Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB

| model | size | params | backend | ngl | n_ubatch | main_gpu | sm | ts | test | t/s |

| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | ---------: | ----: | ------------ | ------------: | ---------------: |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | pp4096 | 380.67 ± 6.25 |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | pp8192 | 360.75 ± 5.05 |

| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | tg128 | 26.29 ± 0.08 |

build: 1a2a8604 (4878)

I would like to hear more abouth the flags, since I am just noob in this...

thank you for your attention to this matter

submitted by /u/OkBase5453
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA