DeepSeek-V4-Flash-0731 unsloth gguf on A100
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
A100 with 40gb VRAM:
- 162GB Q8_K_XL
- ~16.1 tok/s generation
- Only 15.8GB of 40GB VRAM used with all experts on CPU
NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the single 40GB A100 at 17.7 tok/s with 6 experts loaded into VRAM, with Codex driving it through a full agentic coding loop
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.