2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
What I have:
- CPU: EPYC 7551 (32c/64T, Zen 1)
- Board: Supermicro H11SSL-i (SP3), Rev 2.0
- RAM: 128 GB DDR4-2133 (all 8 channels full)
- GPU: 2x RTX 3090 (48 GB total, PCIe 3.0)
- 1500 W PSU
What I run:
- Qwen3-Flash-Next (177B total / ~6B active MoE, IQ4_XS) on Ilama.cpp. Experts live in system RAM, hot ones cached in VRAM. Single stream = 38 tok/s. Two parallel requests drop to ~4 tok/s each.
Budget:
~$800. Realistically that's either one more RTX 3090 or a CPU upgrade (a Zen 2 "Rome" EPYC drops into the same board). A new motherboard is out of budget i think for now.
Which gives more inference speed for this setup - adding the 3rd 3090, or swapping to a faster/newer CPU?
And would more/faster RAM matter here? Curious what people running similar rigs have actually measured.
I am also interested in having multiple agents running at the same time, which currently slows it down heavily, so keeping the performance at multiple agents parallel would be a huge boost as well!
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.