r/LocalLLaMA · · 1 min read

Optimal Configuration for 4x3090s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Optimal Configuration for 4x3090s

Aiming for RTX 6000 like performance at 25% of the cost.

https://preview.redd.it/mi6fqpdd5hih1.png?width=631&format=png&auto=webp&s=9a639bcc79c7eb0a3834a88317220c289b7a52b6

The top 2x3090s are connected via tensor parallelism, then those are connected in a pipeline feeding into another pair that are also using tensor parallelism. The reason being that I don't see anyone getting speedups by putting all 4x3090s in tensor parallelism (actually slower in most cases).

I don't want to have to buy a whole new motherboard for this setup. Currently I have an Asus Proart B850 Creator and Ryzen 7600X CPU. So, I am likely going to purchase a dedicated PCI switch to get the required number of PCI lanes. i.e. Something like this

I am able to fabricate my own brackets and parts now for securing the GPUs in the case. I am absolutely not going to go the open air mining style rig. I want them to all fit in the case securely. (The case is large enough).

Question for the community:

Has anyone else run this configuration before? What kind of inference speed did you get by moving from 2 cards to four?

submitted by /u/Civil_Fee_7862
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA