r/LocalLLaMA · · 1 min read

M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds

I've been playing around with various models on the M5 Ultra 256GB 80-core Mac Studio. These are the results over many rounds of agentic inferencing.

I'm happy with the performance. Glad to have the large amount of RAM. But it does feel like the GPU is underpowered for this amount of RAM. I'm wondering if a 512GB unit for AI inference makes sense at all - because the GPU will be the clear bottleneck.

submitted by /u/dreamingwell
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA