r/LocalLLaMA · · 1 min read

Anyone interested in building a harness-only benchmark?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one.

End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks [1] , grouped by underlying models and reasoning efforts. Anyone can contribute results.

The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.

If there is sufficient interest, I will create a discord.

Disclosure: I am the maintainer of a coding agent called Dirac (https://github.com/dirac-run/dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.

[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.

submitted by /u/Comfortable-Rock-498
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA