r/LocalLLaMA · · 1 min read

40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)

daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster

I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20% faster e2e in reality, if at all, but hey, it's free and open! Apache 2.0

submitted by /u/Dany0
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA