gfx906-llama-cpp: New PP/TG gains for MI50/MI60/Radeon VII/AMD GCN
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Time for another update! We have been busy and managed to improve the gains substantially (mostly from exploring existing llama cpp PRs and adopting relevant things).
Among other things the README.md was also appended to provide a better overall picture of what’s in the fork, why and from whom.
| metric | upstream t/s | fork t/s | gain |
|---|---|---|---|
| prefill PP16384 | 332.5 | ~410 | +23% |
| 120k deep fill | 231.4 | ~264 | +14% |
| TG | 13.6 | ~15.1 | +11% (parity pre-mirror) |
| context | cannot fit | 250k on 40 GB | tight-fit machinery |
| outputs | - | - | bit-identical (sha + token-for-token) |
https://github.com/milpster/gfx906-llama-cpp/blob/master/README.md
(Yes i made this with AI)
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.