r/LocalLLaMA · · 1 min read

366 t/s Qwen3.6 27B NVFP4 on v100s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

These are single stream numbers

Following on from my previous post about v100s (here) and inspired by this comment (here) I decided to work on kernels that allow for an extremely fast path for Nvfp4 weights on sm70 and almost free deep speculation on sm70 as well.

Which leads me excitedly on to the launch of “v100-skinny” (cause the kernels are skinny)

My work can be found here: https://github.com/dnv2003/v100-skinny

Many caveats about the quoted number in the title are in the repo but it is the absolute best case for mtp that being extraction. However you can expect around 240 on structured generation like json and 200 on mtp friendly code (think boiler plate,patterns, html etc using the “flagship configuration of k=7”)

submitted by /u/Simple_Library_2700
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA