Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hello guys,
I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageable.
Did someone tried out in some reasonable GPU sizes up to 48GB VRAM?
Thanks for the feedback!
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.