I collected every single LLM coding benchmark, and computed their Intelligence Density
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| The intelligence in my context is an aggregate index, I called the Agentic Coding Index, across most relevant agentic coding benchmarks: SWE-bench Pro, DeepSWE v1.1, Terminal-Bench (v4, v3, v2.1), Code Arena Elo, and LiveCodeBench v6. Intelligence/Parameter=Scale x (Agentic Index / Norm) ^ (Super_Linear_Exponent) / sqrt(PCount + PLowerBound)
Agentic Coding Index: DeepSWE v1.1 (20%), Code Arena Elo (20%), Terminal-Bench v4.0 (15%), SWE-bench Pro (15%), Terminal-Bench v3.0 (13%), Terminal-Bench v2.1 (12%), and LiveCodeBench v6 (5%). Data Integrity: All benchmark scores are curated from verified public and official sources (model creators, peer-reviewed evaluation reports). [link] [comments] |
More from r/LocalLLaMA
-
What are some practical tasks I can assign to my local AI models?
Sep 8
-
I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)
Sep 8
-
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin
Sep 8
-
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
Sep 8
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.