r/LocalLLaMA · · 1 min read

For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I've been making a lot of comments about optimal setup for Strix Halo (gfx1150) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device.

Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real:

Note: Official llama.cpp running Qwen38FN at 2xt/s and 2xxt/s prefill - 50% theory.

Hopefully this will be helpful to the Strix Halo users.

submitted by /u/feelspeaceman
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA