r/LocalLLaMA · · 1 min read

16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

GLM-5.2 UD-Q4_K_XL GGUF @ 12.2 tok/s output // 30.9 tok/s input on a real 10.7k-token document using llama.cpp RPC - At 10.7k context: 10.2 tok/s output with coherent long-form generation

  • Two parallel requests: 14.5 tok/s aggregate

Context: 2x 16,384-token slots

  • Model size: 436 GiB

Hardware: 16x AMD MI50 32GB, 512GB total, 100w cap each

VRAM, split across two 8-GPU nodes

  • Interconnect: direct 10 GbE DAC, MTU 9000

Before llama.cpp, spent a lot of time integrating GLM-5.2 AWQ INT4 into a custom vLLM-gfx906 v19 Moby Dick build using TP=8 and PP=2. Hit 14t/s decode and 50 t/s prefill but inference degraded after 10k. Didn't get around to MTP things yet. Here's hoping someone figures out getting the Moby Dick repo going with it

submitted by /u/Legal-Ad-3901
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA