r/LocalLLaMA · · 1 min read

M5 Max users: what models are you using & what tk/s are you getting?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I was using antirez’s ds4 for a while and getting around 20 tk/s, which worked for my purposes. But I know there have been big advancements between Qwen, the DS4 vision model, and GLM.

I’m not sure how the quants affect performance, so what’s the best thing to run right now & how fast is it?

submitted by /u/A_Wild_Entei
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA