How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a benchmarking issue? EDIT: The contribution of this benchmark to the Intelligence index of artificialanalysis.ai: Full Intelligence Index v4.1 weights: GDPval-AA v2: 20% Terminal-Bench 2.1: 16% τ³-Bench Banking: 14% Humanity's Last Exam: 12% AA-Omniscience Accuracy: 8% SciCode: 8% GPQA: 6% AA-LCR: 6% CritPt: 6% AA-Omniscience Non-Hallucination: 4% Source: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1 [link] [comments] |
More from r/LocalLLaMA
-
Luth-2: New State-of-the-Art French Small Language Models
Aug 11
-
I ran Muse Glimmer @ 1M context - All tests passed.
Aug 11
-
Qwen 3.8-27b coming this week
Aug 11
-
Nvidia reportedly testing lower memory configs of Rubin Ultra as memory shortage bites back — designs tested include as little as 192 GB and step back to HBM4
Aug 11
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.