MoE models around A2B
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources.
What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any.
- LFM2 24B A2B (5 months old)
- Mellum 2 12B A2.5B (2 months old)
- Moondream 3.1 9B A2B (This month)
- VAETKI 20B A2B (7 months old) (Talk about an unknown model, it has one mention on this sub)
- DeepSeek V2 Lite 16B A2.4B (2024. Remember when DeepSeek was making SMALL models?)
- Ring Mini / Ling Mini, 16B A1.4B (2025)
- There are also • at least three • Nemotron fine tunes 12B A2B, and a 23B A2.8B, some 1-2 months old. Not sure what the deal is with those.
Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects.
Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down?
[link] [comments]
More from r/LocalLLaMA
-
35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder
Aug 25
-
Granite Speech 5.0 Turbo CTC: Extremely Fast and Accurate Transcription
Aug 25
-
Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀
Aug 25
-
Mac Studio M5 Max Cost Analysis
Aug 25
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.