Anyone else completely tuning out these massive "open weight" drops?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Tbh the benchmarks on stuff like GLM-5.2 look insane. 753B params, 1M context, MIT license... everyone is throwing a party on the front page right now. But like... what is actually "local" about this anymore? A 700B+ MoE isn't fitting on anyone's home rig. Even if you absolutely crush it down to a q1 or q2 GGUF, you're still not running it. I've got a dual-GPU setup running on a solid x8/x8 bifurcation board, and even heavily optimized under AMD ROCm, these behemoths are physically impossible to load without an enterprise server rack. I miss when this sub was actually about self-hosting. The vibe used to be sharing compile tricks, fighting with llama.cpp --batch sizes, testing new quants, and actually squeezing models into our hardware. Now half the posts are basically just free marketing for models that 99% of us can only use by paying for APIs or renting cloud instances. Which completely defeats the purpose. Don't get me wrong, it's cool that the weights are actually published instead of locked in a vault. But practically speaking? They might as well be closed source for normal people. Maybe I'm just salty about being VRAM poor lol, but the hype for these giant unrunnable drops is totally dead for me. Anyone else feeling this?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.