MiMo v2.5 is underrated. Feels like the tokens are pouring out of the screen in OpenCode.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| When I recently built my inference server, I expected to deploy DeepSeek v4 flash, but that doesn't look like it's going to be fast for a long time, if ever. There is a massive gap, as we all know, in competent models between 30b and 400b. I was very surprised to find that this is the best model by far that falls within this gap. MiMo v2.5 the only local model I have seen that's faster than fast cloud providers. It is actually worth it to build a big server just to run this thing. Running 192GB 4090 VRAM, I tested Bartowski IQ4_XS, IQ4_NL, Unsloth UD-Q4_K_S, and gghfez "unfused" IQ4_XS in ik_llama.
Looping
Optimizations
Multimodal
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.