any one else finds Mimo v2.5 better than deepseek v4 flash!?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I noticed while using both, mimo was often better,
after benchmarking mimo v2.5 via open code endpoint in diff harness like codex, oh my pi, hermes. i found that mimo is indeed better in coding tasks.
and over all, hermes scored 55% with mimo v2.5 via terminal bench v2.0
others did under 50% too with any harness from my list or deepseek v4 flash
Not that i dont like deep seek v4 flash its GOAT, i have used it more. but as per benchmark both models are same at most places but when u run real life complex problems solving mimo v2.5 seemed to me helping me more
i tested hy3 preview too. idk to me it felt like benchmark trained. needs to try more , but scores were pretty low for me in terminal bench
EDIT: also while comparing oh my pi vs hermes vs codex cli. found hermes better for some reason. (offc for low lvl models only in my casestudy)
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.