DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hello everyone I want to join the hype of posting specs.
CPU: AMD EPYC 74F3 24-Core
RAM: 8 Channel 3200 DDR4
GPU: RTX A6000 48GB
Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k context). Inference is a steady 17.20t/s~ and the 48GB VRAM is enough to have the full 1mil context but PP will be so bad. Sadly not as cool like those M5 Macs.
Anyone else having similar specs?
Edit: I was informed about batch size and set mine to 8096 and my Prompt processing jumped to almost 400t/s at the start. it got to around 300t/s at 20k context. Better than my 70t/s stock lol.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.