r/LocalLLaMA · · 1 min read

Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

sudoingX on X spent a week on this and posted his wrap-up. The hardware is his. He isn't affiliated with us and we didn't see any of it before he put it out — I work on Ling at inclusionAI.

Where he landed on one Spark: 38.7 tok/s on the official INT4 once it's configured right, 35.2 on the community GGUF, 2.4x what DeepSeek V4 Flash does on the same machine. His phrasing is that Ling-3.0-flash earned a permanent seat on his box.

What got my attention was the correction in the middle. He'd posted that our official quants don't run on a single Spark. Two days later he posted that they do, that they're the fastest path he has now, and quote-tweeted his own earlier post to say so.

Most of those numbers have the method sitting next to them on his timeline, so you don't have to take it from me.

One Spark and one person, over a week. He's been at this longer than I've been reading about it.

submitted by /u/AcanthisittaOk1699
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA