r/LocalLLaMA · · 1 min read

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization.

No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some noticeable changes:

  • Flash Next is still king, but the benefit of xhigh vs medium reasoning effort is now properly showing.
  • Same for 3.8 27B (however in everyday tasks I personally still prefer using medium)
submitted by /u/WonderRico
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA