r/LocalLLaMA · · 1 min read

1x32GB V100 vs 2x16GB V100 vs 5060ti 16GB for QWEN 3.8

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hi All,

I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux.

Objective is to increase context and use a better quant and also free up 5060 for other tasks.

The options I am considering are 1x32GB V100 and 2x16GB V100.

Theoretically, 2x16GB should be superior in terms of performance to 5060 and 1x32gb due to higher memory bandwidth.

One issue I have to deal with is that I am limited in terms of CPU to GPU comms - I only have 2x x4 lines available.

Any other good options in the same price range?

UPDATE: Found a very interesting page, showing performance of multiple V100 with qwen 3.8 : https://domoticx.net/docs/llm-with-lama.cpp

submitted by /u/ColorsOfCosmos
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA