r/LocalLLaMA · · 1 min read

One more 'you should try ExllamaV3/exl3 for flash next' appreciation post

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

After seeing a few posts on here about it, I finally tried exl3 3bpw and exllamav3 for running flash next - with amazing results.

On 3x3090s, 128GB DDR4: 1500 prefill, 80 tps decode

On 1x5090, 128G. DDR4: 1500 prefill, 29 tps decode

Both at 262k context, both with vision/spec decoding. Really impressed, definitely replacing vllm/llama.cpp for me on this model.

Quant capacity seems good so far, going to gest the 4bpw later for comparison.

Check it out if you were sleeping on it like I was!

submitted by /u/youcloudsofdoom
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA