r/LocalLLaMA · · 1 min read

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

A100 with 40gb VRAM:

  • 162GB Q8_K_XL
  • ~16.1 tok/s generation
  • Only 15.8GB of 40GB VRAM used with all experts on CPU

NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the single 40GB A100 at 17.7 tok/s with 6 experts loaded into VRAM, with Codex driving it through a full agentic coding loop

submitted by /u/Different-Pickle1021
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA