r/LocalLLaMA · · 1 min read

How to even compare quants from various sources?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

How do you guys deal with so many variables? So many publishers, each calling their quants "best", and then it's a mess to manage, download the weights, tweak the temperature etc. for each source?

It is relatively simple if I'm comparing different quantization levels (like Q4 vs Q5), that's mostly linear and I run the best one I can afford.

Also how do you even arrive at right parameters to use for each (temperature, top-k, penalty.. the whole bunch) and how to get an overall best? Do you just leave them at default? This is like a 20 dimensional optimization problem, except each eval takes SOO LONG (download, run, configure etc.)
I spent hours comparing but couldn't really conclude anything. Super confused. Need help.

submitted by /u/Glad_Claim_6287
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA