r/LocalLLaMA · · 1 min read

Dual DGX Sparks- 40tk/s single 1M ; 350 tk/s agg. - Deepseek V4 Flash (vs RTX Pro 6000 vs Mac M2 Ultra 192)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

First of all shout out to Aiden/Antirez & geniuses at the Nvidia community threads. I'm merely claude-vibing off of their works.

That a said, i thought i'd share recipes & learnings & benchmarks so far on running big MOE models on two dgx sparks at a reasonable speed for agent use:

https://github.com/elsung/dgx-spark-deepseek-v4-flash

The kicker here is that you need 2 DGX sparks to really get the speed we need, and you have to spend the $180 on that single cable for 200G/s over connectx7 in order to get this speed.

BUT, being able to run ~40tk/s on a model that is arguably in the same playpen as the frontiers is exciting and something myself and others probably have been striving/dreaming about for some time now.

I also put in benchmarks against the RTX Pro 6000 and the Mac M2 Ultra 192GB.

TLDR;
- Dual DGX - FP8 ~40 tk/s
- Single DGX - FP8 ~14 tk/s
- RTX Pro 6000 - Q2 ~46 tk/s
- M2 Ultra 192GB - Q2 ~29 tk/s

2x DGX wins cuz FP8 & fast and can run concurrent.

up to 350 tk/s aggregate running 32 requests at 256k context each.

Hopefully this is useful for other folks~

Credit links / Threads (ongoing discussions here)

submitted by /u/elsung
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA