We've gotten some great medium sized models lately (DSV4 Flash 0731, Inkling Small, Laguna S 2.1, Step 3.7 Flash) but does anybody else want to see some new 70-80b contenders?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B.
On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is not the worst but it does get a bit annoying on agentic coding tasks.
If we could get some new MoE 70-80B models that are smarter than the Qwen 3.6 family, I would be so happy.
Double-ish the prefill/gen would make all the difference.
Maybe this is my fault for being cheap and building my GPU rig with some V620's but that price-to-VRAM ratio is hard to beat and I couldn't justify spending more than that so here we are.
Or does anyone have some tips? I've been using ROCm + llama.cpp -- I tried using -sm tensor to speed things up, but it's slower. And it gets slower and slower as I try to enable more GPUs with it. So I'm just back to layer split.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.