r/LocalLLaMA · · 1 min read

32 total local models tested head to head

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I ran 32 local models head to head on one fact-extraction corpus, 1,001 notes, paired bootstrap on every adjacent pair. Several weeks of compute time, all on consumer grade cards.

Most of the field does not separate. Six consecutive steps from 2B to 31B, and the bootstrap cannot order a single adjacent pair. The top two do not separate from each other either: a 35B MoE against a dense 27B from the same family, -0.0106, CI [-0.0294, +0.0088].

LFM2.5 is the exception, in the wrong direction. It landed two days ago and loses to models a fraction of its size. LFM2.5-8B-A1B scores 0.5198 and LFM2.5-2.6B 0.5854, against 0.6406 for gemma-4-E2B at 2B. E2B's worst quant still scores 0.6017, ahead of both

https://rakuensoftware.com/blog/local-llm-fact-extraction-head-to-head

submitted by /u/KitchenAmoeba4438
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA