Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| The interesting part of this one isn't the top number, it's how little distance there is between the top and the bottom of the ladder. Where it comes from: I work on Ling at inclusionAI, these aren't my numbers. sudoingX on X benched the full community GGUF ladder on his own DGX Spark, posting his results with permission. Single stream decode: Q5_K_M, 40.2 tok/s, fastest and near-lossless Q4_K_M, 38.2 tok/s, smallest footprint Q6_K, 32.0 tok/s, max quality for about 16% off the top 32 to 40 across the whole ladder. With 5.1B active out of 124B, so few params fire per token that the quant barely moves decode speed. That's not how this goes on a dense model, where dropping bit width usually buys you real throughput. Q5 landing as both the fastest and the near-lossless pick is the useful part. Normally that's a trade and you have to decide which one you care about. Here the sweet spot isn't a compromise, it's just the answer. For scale on the same box, he measured DeepSeek V4 Flash at 16.5 tok/s, so Q5 is about 2.4x that, and even max-quality Q6 is close to 2x. Charts are his. If anyone has a Spark and gets a different curve, post it. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.