I'm really hoping we're in 2026's 2-month-gap between QwQ and Qwen3 right now
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every possible word, I'll notice the right one!") kinda made it unusable for agentic coding.
It was ~2 months later that Qwen3-32B came out which delivered QwQ's peaks with usable amounts of reasoning.
I know some people are having a great time with Qwen3.8-27B, and same, but I can't have a good sit-down session with it because the reasoning takes so damn long. Everything I do with it needs to be async or compromise on quality (it's still great when you limit reasoning but definitely loses that next-gen edge). I also have to watch context like a hawk.
Maybe 3.8 is 2026's QwQ and a competitive model requiring less reasoning is just around the corner?
[link] [comments]
More from r/LocalLLaMA
-
Apple releases M5 ultra at 1.2TB/s bandwith
Aug 25
-
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
Aug 25
-
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Aug 25
-
Qwen 3.8 Flash Next day 0 support from unsloth
Aug 25
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.