AMD Lucebox Beats Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey fellow llamas, sorry for posting again this week but i thought this was interesting to showcase to share with y'all. Lucebox partnered up with AMD to bring heterogenous consumer hardware to life. We worked really hard on this, and were able to have Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) to beat one NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed. R9700 takes the dense path, the hot experts, the cache and the draft model. Strix Halo holds the other experts in 128gb and computes them at the same time, not after. 51.1 tok/s on the full 284B, 3.63x one DGX Spark, and $2,899 less than having two of them. You can check the full technical breakdown here: https://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism Please don't be rude with us, we know that it's still something experimental running q2 model at 16k context. We're currently working on implementing KVFlash to get to context at 64k-128k. Let us know what you think and if you have any feedback! [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.