Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running\_qwen3\_30b\_a3b\_at\_50\_toks\_on\_rtx\_5060\_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta Network kernels later and here it is. It runs at 55 tok/s (61 when not recording - since the recording eats cpu and gpu capacity), significantly outperfroming llama.cpp running qwen3.5 35B at Q8 quant. Mind you - this is without MTP. MTP can speed up the generation even further, for a few cool reasons beyond the basic multiple tokens produced. Working on a blog post documenting the trick used to get this high speed up. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.