Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| mlx-dspark is an MLX port of DeepSeek's DSpark speculative-decoding drafters (the DeepSpec release), plus z-lab's DFlash, with one lossless verify loop. v0.10.0 adds Qwen3.8-27B via RadixArk's drafter, the first SpecForge/SGLang-packaged head it loads. Numbers (M4 Pro 48 GB, medians of 3, greedy, output ids identical to plain decoding):
"Lossless" is checked, not asserted: the target verifies every drafted token, and the Mac app's Race view runs speculative vs plain on the same prompt and diffs the token ids (video is that view). Everything is Repo: github.com/ARahim3/mlx-dspark I'd appreciate any feedback you might have after using it. [link] [comments] |
More from r/LocalLLaMA
-
Any upcoming models to be excited about?
Aug 23
-
you can now use MTP in GLM-Air
Aug 23
-
Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system
Aug 23
-
We quantized Qwen 3.8 27B and compared the quants on an RTX 6000
Aug 23
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.