r/LocalLLaMA · · 1 min read

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Extened garlic to run Qwen3.5 35B A3B float8 at 55 tok/s on RTX 5060 Ti

In a previous post (https://www.reddit.com/r/LocalLLaMA/comments/1utefpr/running\_qwen3\_30b\_a3b\_at\_50\_toks\_on\_rtx\_5060\_ti/) there seemed to be great demand for bringing in Qwen3.5 35B. Some Gated Delta Network kernels later and here it is. It runs at 55 tok/s (61 when not recording - since the recording eats cpu and gpu capacity), significantly outperfroming llama.cpp running qwen3.5 35B at Q8 quant.

Mind you - this is without MTP. MTP can speed up the generation even further, for a few cool reasons beyond the basic multiple tokens produced.

Working on a blog post documenting the trick used to get this high speed up.

submitted by /u/Azazelionide
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA