r/LocalLLaMA · June 14, 2026 · 1 min read

Dual DGX Sparks- 40tk/s single 1M ; 350 tk/s agg. - Deepseek V4 Flash (vs RTX Pro 6000 vs Mac M2 Ultra 192)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

First of all shout out to Aiden/Antirez & geniuses at the Nvidia community threads. I'm merely claude-vibing off of their works.

That a said, i thought i'd share recipes & learnings & benchmarks so far on running big MOE models on two dgx sparks at a reasonable speed for agent use:

https://github.com/elsung/dgx-spark-deepseek-v4-flash

The kicker here is that you need 2 DGX sparks to really get the speed we need, and you have to spend the $180 on that single cable for 200G/s over connectx7 in order to get this speed.

BUT, being able to run ~40tk/s on a model that is arguably in the same playpen as the frontiers is exciting and something myself and others probably have been striving/dreaming about for some time now.

I also put in benchmarks against the RTX Pro 6000 and the Mac M2 Ultra 192GB.

TLDR;
- Dual DGX - FP8 ~40 tk/s
- Single DGX - FP8 ~14 tk/s
- RTX Pro 6000 - Q2 ~46 tk/s
- M2 Ultra 192GB - Q2 ~29 tk/s

2x DGX wins cuz FP8 & fast and can run concurrent.

up to 350 tk/s aggregate running 32 requests at 256k context each.

Hopefully this is useful for other folks~

Credit links / Threads (ongoing discussions here)

Antirez & his awesome work
- https://github.com/antirez/ds4
Aiden thread & DGX threads i found via Nvidia Communty threads:
- https://forums.developer.nvidia.com/t/deepseek-v4-flash-aiden-recipe-from-reddit-1m-token-session-operational-cuda-12-1-tailored-for-dgx-spark-gb10/372268/61
- https://forums.developer.nvidia.com/t/deepseek-v4-flash-official-fp8-running-across-2x-dgx-spark-tp-2-mtp-200k-ctx-recipe-numbers/370309

submitted by /u/elsung
[link] [comments]

Discussion (0)

No comments yet. Sign in and be the first to say something.

Discussion (0)

More from r/LocalLLaMA