Where are the current GPU VRAM sweet spots?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I have been reasonably satisfied with my single R9700 (32GB) as I can run practical quants of Qwen 3.8-27B at good speeds, as well as other similar models in its weight class (Gemma 4 is still my go-to for general knowledge, until I see something better - has that happened?).
But my inference box has a second PCIe x16 slot and while while the motherboard/CPU will only run both slots at x8, that second slot still whispers to me like the Green Goblin mask. R9700s are reasonably affordable still, but with AMD indicating that it's going to jack prices soon, I'm wondering whether it makes sense to pull the trigger on a second card before the price hikes?
Going to two cards probably only means a PSU upgrade, whereas going beyond that will mean motherboard/CPU/RAM upgrade, which is not really practical right now. My searches and discussions with commercial LLMs make me think the newer open-weight releases are trending towards larger MoEs that won't fit in 64GB at an acceptable quant, and that's making me think that a single-card setup is probably the best bang-for-buck when it comes to local inference. Am I missing something?
[link] [comments]
More from r/LocalLLaMA
-
I really don't understand Jev hype
Sep 21
-
Clarification on the Qwen-image-2.1 license
Sep 21
-
What's the verdict on Ternary Bonsai 2 27B?
Sep 21
-
mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.
Sep 21
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.