r/LocalLLaMA · · 2 min read

Add a amd 9700 ai pro to a 3x5090 system vs buy a 5070ti for general useage.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hey guys I have this OCD im trying to decide about, I was lucky enough to buy 3 5090 before all the crazy ai stuff started and while that system works fine. Id also bought a razor core egpu that stopped working a little ago and so got sent off for repair, well its come back now and I'm debating if to move one of those 5090s back into the egpu and now have a long pci 4 risor free that could add another gpu to system (ive tested 4 gpus on the proart 3 5090 + a 4090 (in egpu) i stupidly sold to recover some of cost, so I know it works. Yes I know im lucky and spend way too much for a hobby but its what I enjoy doing, I wanted to build a epic system but I know thats out with price of ram.

I know about the Cuda + Rom mixing issue but know about using the RPC server with loopback trick for a mixed system, so I wouldn't have to use vulkan, but id be limited then to having to use llama.cpp. I'd get 128 gig vram, so my dream would to be able to run the new DeepSeek I'm obsessing about. (i can run it with offload to cpu but its less than 100/t/s prefill.

Or the otherhand i could just get a cheaper gaming gpu like a 5070ti, id have less vram but leave those 3 5090 for AI stuffs. Is Vram king? Will adding a much slower amd card to mix sorta make it pointless to have 128 gig vram?

I also have a second machine that used to use with RPC / one of these 5090 before i got the proart. It has 2 x pci 5x8 slots i could maybe add something like 2 5060ti for cost of above options. But from my testing RPC with moe models seems to generally suck. Only ever had good speedups when using dense models.

Anyone have any suggestions or ideas before i go drive over tomorrow to buy a new psu for the egpu lol. Thank you.

submitted by /u/fluffywuffie90210
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA