16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
GLM-5.2 UD-Q4_K_XL GGUF @ 12.2 tok/s output // 30.9 tok/s input on a real 10.7k-token document using llama.cpp RPC - At 10.7k context: 10.2 tok/s output with coherent long-form generation
- Two parallel requests: 14.5 tok/s aggregate
Context: 2x 16,384-token slots
- Model size: 436 GiB
Hardware: 16x AMD MI50 32GB, 512GB total, 100w cap each
VRAM, split across two 8-GPU nodes
- Interconnect: direct 10 GbE DAC, MTU 9000
Before llama.cpp, spent a lot of time integrating GLM-5.2 AWQ INT4 into a custom vLLM-gfx906 v19 Moby Dick build using TP=8 and PP=2. Hit 14t/s decode and 50 t/s prefill but inference degraded after 10k. Didn't get around to MTP things yet. Here's hoping someone figures out getting the Moby Dick repo going with it
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.