r/LocalLLaMA · · 1 min read

Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP

https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139

I've only recently started looking deeper into vLLM after running llama.cpp for a good while.

Initially vLLM (official repo) was terribly slow on my four R9700s (tried that one with two as well), however after trying radiance everything changed.

Prefill 17636 - TG at that time was 36,6

That prefill spike was two agent profiles working on different tasks simultaneously (one is writing a yt-dlp dl/conversion workflow the other is auditing agents (profiles).

Best TG i've hit was 106 Tok/s with a 80% MTP 4 acceptance rate.

For reference, I'm running a Gigabyte MZ32-AR0 (Rev 1.0), EPYC 7282 and using Hermes with Qwen 3.8 27b fp8 262k ctx - worth noting that one GPU is actually only running by PCIe 4x8, three full 4x16.

On that note i'm also happy to say that vLLM-Radiance does work well with a quad setup in my case - nvtop consistently shows 100% usage of the four cards, officially only dual setups are supported.

I hope this doesn't count as a low effort post, i just had to share.

//E

submitted by /u/im_EDEN
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA