r/LocalLLaMA · · 1 min read

Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hi all, I'm looking for the best model for a hobby project and trying to make sense of the various data I came across. I know benchmarks do not often translate to the real world, especially to your particular use case (whatever it may be). But this is truly baffling:

AA says Qwen 3.8 27B is better by miles: https://artificialanalysis.ai/models/comparisons/qwen3-8-27b-vs-gemma-4-31b?intelligence-comparison=intelligence-vs-end-to-end-response-time

While Arena says Gemma 4 31B is almost 20 places ahead and completely trounces Qwen in many categories: https://arena.ai/leaderboard/text/overall

The sentiment in this sub definitely seems in favour of Qwen, although not necessarily against Gemma which I think is still considered a good model. I recall poeple saying Qwen tends to be more tenacious and better at reasoning although at the cost of overthinking simple things.

What is your explanation or experience with these models?

submitted by /u/uncle_leon
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA