r/LocalLLaMA · · 1 min read

[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=webp&s=5faeec866acc9e8eca8dc9e84a6661bee7bcf954

https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=webp&s=440320ca4646ff7dceaf2aa5ac2ef76faeeba525

New Qwen3.8-Flash-Next quantization using GSQ-RCO. Cuts the size of Qwen3.8 Flash Next from around 80-95GB to 68-76GB, while still preserving near baseline quality. Also their Q2_0 variant claims to be much faster offering 6.2x better prompt throughput in coding.

"Q2_0 is built for speed. It avoids the quantization formats that rely on large lookup tables: those formats pack more accuracy into a given bit-width, but decoding them costs real time, and on this model that cost dominates inference. Q2_0 delivers 3.4x the prompt throughput and 1.9x lower end-to-end latency than IQ2_XS at a slightly smaller file size, and its decode rate stays flat across workloads instead of varying with the content. The trade is a little quality: 89.07 task average against 89.16 for IQ2_XS, and 3.5 points below IQ3_XXS. Pick it when throughput matters most, and see Performance for the measurements.

The IQ3_XXS model is the strongest operating point: it matches the base model exactly on AIME25 (100.00) and is within 0.51 points on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of the BF16 size."

Model link: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

submitted by /u/BullfrogScary8947
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA