Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling:
AA says Qwen 3.8 27B is better by miles: https://artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b?intelligence-comparison=intelligence-vs-end-to-end-response-time
While Arena says Gemma 4 31B is almost 20 places ahead and completely trounces Qwen in many categories: https://arena.ai/leaderboard/text/overall
The sentiment in this sub definitely seems in favour of Qwen, although not necessarily against Gemma which I think is still considered a good model. I recall poeple saying Qwen tends to be more tenacious and better at reasoning although at the cost of overthinking simple things.
What is your explanation or experience with these models?
[link] [comments]
More from r/LocalLLaMA
-
Can we reconsider the megathreads?
Aug 26
-
Are models with N-Gram tables going to completely change the AI race?
Aug 26
-
Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you
Aug 26
-
[Megathread] GLM-5.3-Flash - former ox-alpha
Aug 26
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.