I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8...
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
TL;DR: I run a mismatched Tesla V100-PCIE pair—one 16 GB card and one 32 GB card, 48 GB total—in a Proxmox/LXC-based local-inference lab. The practical winner so far is a recent CUDA build of llama.cpp with tensor split, Flash Attention, --numa distribute, and large batches. On Qwen3.8 27B Q6_K_M I measured 1,376.9 prompt tok/s at 2k, 1,324.3 prompt tok/s at 4k, 1,221.5 prompt tok/s at 16k, and 39.9 decode tok/s. I also made Qwen3.8 Flash Next Q4 GGUF of a 177 B/6B-active MoE model run across the same two cards at roughly 26 tok/s decode and 380 PP.
This is not a claim that old V100s beat modern GPUs. It is a report of what they can do in this crazy "a kidney for a GPU" market!
My Hardware:
| Layer | Hardware / configuration | |---|---| | GPU 0 | NVIDIA Tesla V100-PCIE, 16,144 MiB VRAM, compute capability 7.0 | | GPU1 | NVIDIA Tesla V100-PCIE, 32,494 MiB VRAM, compute capability 7.0 | | Server topology | NUMA-aware host; benchmarks use `numactl --interleave=all` or `--numa distribute` | | Virtualization | Proxmox with LXC inference containers for isolated runtimes | | CPU| 2 x Xeon E5-2696 v4 @ 2.20GHz| | RAM | 512 GB DDR4 2400 ECC RDIMM| | CUDA stack | CUDA 12.8; NVIDIA driver libraries are host-mounted into the vLLM LXC where needed | | Workload | Runtime / model | Best result | Configuration behind the result |
|---|---|---|---|
| Best 27B prefill | Mainline llama.cpp, Qwen3.8 27B Q6_K_M | 1,376.87 pp tok/s at 2k | Tensor split; main GPU = 32 GB V100; Flash Attention; Q8 KV; batch/ubatch 2,048; 20 threads; NUMA distribute |
| Best 27B decode | Mainline llama.cpp, Qwen3.8 27B Q6_K_M | 39.88 tg tok/s | Same configuration |
| 27B at 16k context | Mainline llama.cpp, Qwen3.8 27B Q6_K_M | 1,221.49 pp tok/s | Same configuration |
| Best Q8 27B prefill | Mainline llama.cpp, Qwen3.8 27B Q8_K_XL | 1,059.67 pp tok/s at 2k | Tensor split; long-context run used 1:1.5 split |
| Q8 27B at 64k | Mainline llama.cpp, Qwen3.8 27B Q8_K_XL | 664.36 pp tok/s | 65,536-token prompt; tensor split 1:1.5 |
| Best ik_llama 27B result | ik_llama.cpp, Qwen3.8 27B Q6_K_M | 670.61 pp tok/s at 2k; 35.43 tg tok/s | Graph split; 1:1 split; smaller batch settings than the mainline winner |
| Largest model tested | ik_llama.cpp, Qwen3.8 Flash Next Q4_K_XL | 380.67 pp tok/s at 4k; 26.29 tg tok/s | 103.68 GiB GGUF; reported 177B total / 6B active MoE; graph split 1:1.75 |
| ninfer-v100 | NInfer, Qwen3.8 27B NVFP4 | 927.54 pp tok/s at 8,190 | 19.7 GiB weights; INT8 KV; MTP |
What the tuning showed
| Qwen3.8 27B Q6_K_M, mainline llama.cpp | pp2048 | tg128 |
|---|---|---|
| Layer split, 79 threads | 981.79 | 22.88 |
| Tensor split, 20 threads | 1,371.57 | 38.79 |
| Tensor split, main GPU = 32 GB V100 | 1,376.87 | 39.88 |
| Tensor split 1:2 | 980.71 | 26.10 |
| No split | 967.36 | 18.05 |
Takeaway: on this uneven V100 pair, tensor split—not layer split—and using the 32 GB V100 as main GPU are decisive. The best daily-driver path is mainline llama.cpp + Qwen3.8 27B Q6 + tensor split + Flash Attention + Q8 KV + NUMA distribution.
root@llama-cpp:~# numactl --interleave=all /opt/llama.cpp/build/bin/llama-bench -m /path/to/Qwen3.8-27B-UD-Q6_K_M.gguf -ngl 999 -sm tensor -mg 1 -fa on -b 2048 -ub 2048 -t 20 -p 512,2048,4096,16384 -n 128 -r
5 -lm none --numa distribute
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 48638 MiB):
Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB
Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB
| model | size | params | backend | ngl | threads | n_ubatch | main_gpu | sm | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | -------: | ---------: | -----: | --: | ---------: | --------------: | -------------------: |
| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp512 | 1156.94 ± 5.10 |
| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp2048 | 1376.87 ± 4.59 |
| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp4096 | 1324.34 ± 7.37 |
| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | pp16384 | 1221.49 ± 3.87 |
| qwen35 27B Q6_K | 21.49 GiB | 27.32 B | CUDA | 999 | 20 | 2048 | 1 | tensor | 1 | none | tg128 | 39.88 ± 0.04 |
build: b29c606 (1)
root@llama-cpp:~# for MAIN in 0 1; do
echo
echo "===== MAIN GPU ${MAIN} ====="
/usr/bin/numactl --interleave=all \
/opt/ik_llama.cpp/build/bin/llama-bench \
-m /path/to/models/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 -sm graph -ts 1/1.75 -mg "$MAIN" --fit 1 --fit-margin 1024 -fa 1 -p 4096,8192 -n 128 -b 2048 -ub 1024 -r 10 -o md
done 2>&1 | tee /root/bench-qwen38-main-gpu-final.md
===== MAIN GPU 0 =====
ggml_cuda_init: found 2 CUDA devices:
Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB
Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB
| model | size | params | backend | ngl | n_ubatch | sm | ts | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | ----: | ------------ | ------------: | ---------------: |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | pp4096 | 356.20 ± 7.68 |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | pp8192 | 361.00 ± 4.72 |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | graph | 1.00/1.75 | tg128 | 26.44 ± 0.11
===== MAIN GPU 1 =====
ggml_cuda_init: found 2 CUDA devices:
Device 0: Tesla V100-PCIE-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB
Device 1: Tesla V100-PCIE-32GB, compute capability 7.0, VMM: yes, VRAM: 32494 MiB
| model | size | params | backend | ngl | n_ubatch | main_gpu | sm | ts | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | ---------: | ----: | ------------ | ------------: | ---------------: |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | pp4096 | 380.67 ± 6.25 |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | pp8192 | 360.75 ± 5.05 |
| qwen4exp 125B.A6B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA | 999 | 1024 | 1 | graph | 1.00/1.75 | tg128 | 26.29 ± 0.08 |
build: 1a2a8604 (4878)
I would like to hear more abouth the flags, since I am just noob in this...
thank you for your attention to this matter
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.