r/LocalLLaMA · · 1 min read

VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hello!

I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment)

I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM.

However seeing that this motherboard supports bifurcation on each slot and I can get 8 gpus at PCIe 4x8 makes me think if this would be a viable upgrade in the future.

I see conflicting info about what the performance results will be. If I understand correctly getting beyond 4 GPUs will drastically hurt my token generation speeds because of the PCIe bottleneck? But is that regardless of what GPUs I'm running?

I know for example RTX 3090 needs more PCIE bandwidth because it's much more performant and will spit out much more data that needs to be synced (pardon my lack of terminology), does that mean that I will have smaller performance penalty from going from 4 to 8 video cards with the 3060s?

Can someone guesstimate what should I expect, right now I get 25 tps with Qwen 3.8 27b Q6, running with llama.cpp in layered mode (three 3060s). I expect VLLM with four gpus will be an upgrade (perhaps I could hit 50 tps?), but what about 8 GPUs?

Will it be lower than my current baseline?

Sorry if I'm being ignorant, I'm kinda new to this and I don't trust chatbots. My mind is set to having a good enough local AI server and I'm trying to get the best bang for my buck and current hardware.

submitted by /u/snakeat3rr
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA