Benchmarked every spec-decode method on Qwen3.6-27B across vLLM and SGLang (single RTX PRO 6000 Max-Q)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Spent the last few days measuring speculative decoding on Qwen3.6-27B (dense, NVFP4) on one RTX PRO 6000 Max-Q, comparing vLLM and SGLang across MTP, DFlash, EAGLE3 and ngram. Same pinned client for every engine and 3 restart-samples per point, so the numbers should be comparable. Spec-Bench, greedy, batch 1, averaged over the 6 categories. Speedup vs each engine's own no-spec baseline:
Two things that ate a whole evening:
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.