GPU Power Efficiency Tips and Tricks
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Ran a quick test sweep across the frequency range of the P100s using Darwin-36B (Qwen 3.6 35B-A3B sort of). Not a comprehensive test. Feel free to share your power efficiency tips, maybe we make a bigger thread as we explore ways to improve our efficiency, not just raw tps. Try it:
Also: capping to 150W (-pl 150) cost us 0.3% on this workload — the caps almost never bind during layer-split serving, so they're free thermal insurance. WTF AM I TALKING ABOUT?: The last 265 MHz costs 53 Watts (enough to light a few rooms) and only buys 16% more speed. That'd be the trade you're refusing. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.