Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey y'all. I'll be concise. TL;DR: DS V4-Flash-0731 @ UD-IQ2_M running fully in VRAM on 3xMI50s (90.9 GB model, 96 GB VRAM). Actual speed on llama-server is: - Text Generation: ~15-16 tokens/second stable. Never dipped below 14 tokens/second, even when the model was spitting out a 30K token long reply. llama-server CLI logs, for those interested: https://pastebin.com/nXy9v0x8 I had a brief conversation with the model. Seemed mostly good. At a glance, I noticed 1 mistake: It mixed up the MI50's memory bandwidth (1 TB/s) with PCIe 4.0's bidirectional bandwidth (64 GB/s). For those interested, I exported the conversation .jsonl from llama-server's web UI. You can find it here: https://pastebin.com/CwHm5cTf I only ran a single coding test, as I don't have too much time to thoroughly evaluate the quality of the quant right now. The test I ran is copied from this post by u/perelmanych from 16 hours ago. Specifically, the rubik's cube test that was shown and coded by DS V4-Flash-0731 through DeepSeek's official API, so I'm guessing it's the full precision model. For a given definition of full precision; it's natively FP4 + FP8 mixed precision. Here is the prompt (same as the one from the aforementioned post) that was used: Here's a pastebin of the HTML code generated by my local DS V4-Flash: https://pastebin.com/43bzF2cm See the attached clip to see it running. I'll refrain from giving my opinion yet on the quality of the local quants because I haven't used it yet to form a well-informed opinion. I'm just, in general, blown away that I can run it locally at all. I do use the DS API frequently as-is, and it's amazing that I have the option of running it locally if I so desire. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.