r/LocalLLaMA · · 1 min read

How do we benefits from 2+ T models?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hey, I’ve been really excited to see the latest models being released, but I keep wondering: what are we actually supposed to do with them?

I have 4× RTX 6000 Max-Q GPUs, 7× RX 7900 XTXs, 5× modded 48GB RTX 4090s, and a lot of DDR5 RAM.... and honestly, I can’t even imagine running Kimi K3 at a genuinely usable speed.

I feel like my setup is already pretty extreme, so I’ve been wondering: what is the point of saying that “local AI is winning” when most users are still running models around the Qwen3.6 range, while even the wealthiest users seem to struggle with slow GLM-5.2 inference?

submitted by /u/zakadit
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA