r/LocalLLaMA · · 1 min read

Minimum VRAM GPU to run DeepSeek-V4-Flash-0731 Q4_K_XL at around 30 t/s ?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hello guys,

I'm curious about running DeepSeek-V4-Flash-0731 locally. Since it’s a Mixture of Experts (MoE) model with only 13B active parameters, I was hoping the VRAM requirements might be manageable.

Did someone tried out in some reasonable GPU sizes up to 48GB VRAM?

Thanks for the feedback!

submitted by /u/Informal-Trouble2183
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA