r/LocalLLaMA · · 1 min read

Getting the most out of MTP

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Getting the most out of MTP

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving a lot of performance on the table if you don't tune n_max. Can be easily missing out on 50-100% of the possible performance on some models.

I ran some benchmarks against the various models I am using on my hardware (p100 + 2xV100) and there are some pretty big differences between model families. Below are some of the results I got. Full details, some other models, including impacts to VRAM and scripts to run the benchmarks are on github here: https://github.com/bradrlaw/ai-server/blob/main/docs/benchmarking.md

Gemma-31b scaled nicely with more n-max

Qwen benefited most from a middle setting

Everyone's favorite scaled well

The smaller Gemma model behaved opposite of the larger one

submitted by /u/bradrlaw
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA