K2 Horizon lineup is out on AA, and once again AA plots are misleading.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| The full K2 Horizon lineup is out on Artificial Analysis. The AA intelligence vs. parameters plots show that - 0.9B and 375B are bad - 3.7B and 7B are SOTA - 36B A4B is SOTA for hardware with poor memory bandwidth (spilled experts, Strix Halo, DGX Spark). I'm going to take the AA Intelligence Index at face value here. This post is not about it. The problem is that these models have a god-awful KV cache design. This means that you really can't use the number of parameters for "best in class" considerations, because these models heavily shift to the right on the plot if you replace parameter count on the X axis with RAM requirements. For Q4_K_M weights, no drafter, no vision, 128k kvarn4 KV cache:
Compare them to
Notes: I don't advise compressing 2~4B models to Q4 and I haven't tested these models' tolerance to weights and kv cache quantization yet. The above choices are just to keep the comparison fair. This awful context design means that
[link] [comments] |
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.