r/LocalLLaMA · · 1 min read

Best current Qwen Flash Next Q4-ish? + worth using?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Im running a 5090 and 64gb of ram, so im limited on what I can run. I have currently been able to fit the following -

Atomic Q4_k_m 4.27bpw @ 31 layers offload

Swift IQ4_xs @ 32 layers offload.

Im about to try the Unsloth IQ4_xs as well.

I could get a "bigger" (non IQ) quant for atomic because its smaller, however they do theirs is obviously different. The unsloth IQ4 is also pretty small, the Swift one is the biggest.

i may be able to jump up one size on something, but it would mean offloading more layers and that would seem to be a significant slowdown. I get around 40tok/s if I stay around the 34-30 range.

any recommendations?

and the next question - I can (and do) also run Q5 and Q6 qwen 27b models. Is the bigger quant of 27b actually going to be more intelligent than the cut-down flash-next builds?

submitted by /u/nirurin
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA