r/LocalLLaMA · · 1 min read

Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. 😵‍💫 Adjusting context will increase/decrease usage on both side.

--tensor-split 499,501
= GPU1 12.5 GB, GPU2 15.4 GB

--tensor-split 501, 499
= GPU1 14.7 GB, GPU2 13.4 GB

Tried --spec-draft-device with CUDA0 and CUDA1 separately, no change at all. (same distribution as above)

Also tried --mmproj-device, no much difference.

Tried --no-mmproj-offload, somehow the lower side get even lower 🫣
= GPU1 14.7 GB, GPU2 12.3 GB

I guess it is related to MTP + Tensor Parallel stuff being concentrated on one GPU. No idea how to solve this.

llama-server \ --batch-size 2048 \ --cache-ram 24384 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --chat-template-file /mnt/AI/models/qwen-chat-template-froggeric-22.5.jinja \ --checkpoint-min-step 1024 \ --ctx-checkpoints 32 \ --ctx-size 192000 \ --fit off \ --gpu-layers all \ --image-min-tokens 1024 \ --load-mode none \ --main-gpu 1 \ --min-p 0.0 \ --mmproj /mnt/AI/models/Qwen3.8-27B-mmproj-BF16.gguf \ --model /mnt/AI/models/Qwen3.8-27B-NVFP4-MID-HIGH.gguf \ --parallel 1 \ --presence-penalty 0.0 \ --repeat-penalty 1.0 \ --spec-draft-n-max 5 \ --spec-draft-n-min 0 \ --spec-draft-ngl all \ --spec-draft-type-k q8_0 \ --spec-draft-type-v q8_0 \ --spec-type draft-mtp \ --split-mode tensor \ --temp 1 \ --tensor-split 499,501 \ --top-k 20 \ --top-p 0.95 \ --n-gpu-layers-draft all \ --no-prefill-assistant \ --reasoning-preserve 
submitted by /u/NickCanCode
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA