[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed. Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16. What's inside
Results: At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.
Coder (capability pruned model): Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits. The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.
Both measured at xhigh reasoning effort. Links
Both repositories ship the complete per-tensor RCO allocation. The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome. From the ISTA Deep Algorithms and Systems Lab. [link] [comments] |
More from r/LocalLLaMA
-
Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)
Sep 30
-
add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp
Sep 30
-
BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B
Sep 30
-
DeepSeek now trained on Ascend 950
Sep 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.