r/LocalLLaMA · · 1 min read

2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

What I have:

- CPU: EPYC 7551 (32c/64T, Zen 1)

- Board: Supermicro H11SSL-i (SP3), Rev 2.0

- RAM: 128 GB DDR4-2133 (all 8 channels full)

- GPU: 2x RTX 3090 (48 GB total, PCIe 3.0)

- 1500 W PSU

What I run:

- Qwen3-Flash-Next (177B total / ~6B active MoE, IQ4_XS) on Ilama.cpp. Experts live in system RAM, hot ones cached in VRAM. Single stream = 38 tok/s. Two parallel requests drop to ~4 tok/s each.

Budget:

~$800. Realistically that's either one more RTX 3090 or a CPU upgrade (a Zen 2 "Rome" EPYC drops into the same board). A new motherboard is out of budget i think for now.

Which gives more inference speed for this setup - adding the 3rd 3090, or swapping to a faster/newer CPU?

And would more/faster RAM matter here? Curious what people running similar rigs have actually measured.

I am also interested in having multiple agents running at the same time, which currently slows it down heavily, so keeping the performance at multiple agents parallel would be a huge boost as well!

submitted by /u/ludos1978
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA