r/LocalLLaMA · · 1 min read

... so, yeah.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

... so, yeah.

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

Dense 3.8-27B is just faster... and maybe better due to quantization level...

EDIT:

Hold a second, Flash-Next is actually performing faster than 27B after some key flags on llama.cpp. it's Holding up to 131K without OOM-ing..... maybe...

0.36.940.283 I srv load: --top-k

0.36.940.283 I srv load: 20

0.36.940.284 I srv load: --ctx-size

0.36.940.284 I srv load: 131072

0.36.940.284 I srv load: --cache-type-k

0.36.940.285 I srv load: q4_0

0.36.940.285 I srv load: --cache-type-v

0.36.940.285 I srv load: q4_0

0.36.940.285 I srv load: --flash-attn

0.36.940.285 I srv load: on

0.36.940.285 I srv load: --load-mode

0.36.940.286 I srv load: mmap

0.36.940.286 I srv load: --lazy-mode

0.36.940.286 I srv load: on

submitted by /u/JLeonsarmiento
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA