r/LocalLLaMA · · 2 min read

I benchmarked IFM/K2-Horizon-7B on 16GB VRAM

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I benchmarked IFM/K2-Horizon-7B on 16GB VRAM

After two days of fighting with benchmarking infrastructure, I finally benchmarked IFM/K2-Horizon-7B on 16GB VRAM. TL/DR: it works, but far behind Qwen-3.8-27B.

Setup

I compared 3 models - Qwen3.8-27B, Ornith-1.5-9B and new IFM/K2-Horizon-7B - on the following setup:

  • All models were fit completely to 16GB VRAM (AMD RX7600XT 16GB, RDNA3), no CPU offload.
  • Weight quants were Q8_0 for Ornith-1.5-9B and K2-Horizon-7B, for Qwen3.8-27B it was GSQ-RCO-IQ3_XXS
  • KV quants were selected from available VRAM at 100k+ context (minimal usable for agentic coding IMO): f16 for Ornith-1.5-9B, Q4_0 for K2-Horizon-7B and Qwen3.8-27B
  • For inference, llama.cpp HIP ROCm build was used
    • Mainline for Ornith-1.5-9B and Qwen3.8-27B
    • MBZUAI-IFM/llama.cpp fork for K2-Horizon-7B
  • LLM server works on Ubuntu-24.04.4, official AMD drivers.
  • For benchmarking, MicroBench-12 benchmark was run under WSL.
    • Benchmark harness: clean pi.dev with pi-sandbox plugin.
    • Completion time limit: 40 min per task

Results

Rank Model and quant Context and KV quant Completed Completed task correctness Wall time (s)
1 Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp 114688 @ Q4_0 15/15 1.000 348.5
2 Ornith-1.5-9B-Q8_0 163840 @ f16 15/15 0.983 184.9
3 K2-Horizon-7B-Q8_0 131072 @ Q4_0 11/15 0.909 498.1

Important notes:

  1. K2-Horizon-7B results were partly contaminated - analysing final results I discovered, that model found and explored other models run results (Qwen and Ornith). I decided not to rerun benchmark since even with this information it failed to score good results. Other model runs were clean.
  2. K2-Horizon-7B thought VERY long, 3 of 4 failed tasks were due to timeout (very generous 40 minutes per tasks though), and 1 task was completed with broken, unverifiable code.
  3. I know that Q4_0 KV is bad (see above why I selected it), but Qwen works fine even with this quant.

Conclusion

  1. Even on 16GB slow 128-bit RX7600XT, Qwen3.8-27B is top, undisputable king - even in IQ3_XSS quant and Q4_0 KV-cache it easily solved ALL the tasks.
  2. Ornith-1.5-9B scores second. Though, it is 2x fast and fully usable for agentic coding.
  3. Unfortunately, new K2-Horizon-7B is last. Maybe, it'd perform better with lesser quantized KV-cache, but 16GB VRAM doesn't provide space for it.

So, no surprise, long live the King Qwen!

submitted by /u/Barni275
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA