1x32GB V100 vs 2x16GB V100 vs 5060ti 16GB for QWEN 3.8
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hi All,
I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux.
Objective is to increase context and use a better quant and also free up 5060 for other tasks.
The options I am considering are 1x32GB V100 and 2x16GB V100.
Theoretically, 2x16GB should be superior in terms of performance to 5060 and 1x32gb due to higher memory bandwidth.
One issue I have to deal with is that I am limited in terms of CPU to GPU comms - I only have 2x x4 lines available.
Any other good options in the same price range?
UPDATE: Found a very interesting page, showing performance of multiple V100 with qwen 3.8 : https://domoticx.net/docs/llm-with-lama.cpp
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.