Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| (Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!) About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading. Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload Well, quick confession first. I actually shelved that project shortly after. The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence. Didn't really like that tradeoff, so I abandoned it and never released it. Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops. I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners. My setup
And if you've tried running a 68GB MoE on a 16GB RAM machine with stock I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around. Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using So I went back to my original idea and started implementing it properly as an optional feature inside
The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock. The numbersAll tests below were run with Model: Qwen 3.8 Flash Next IQ2_XS (68GB) Hardware: RTX 3060 12GB + 16GB DDR4 RAM
The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency. The prefetching version is where things get interesting. Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock. That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project. What's still roughIt's not all perfect yet. There are a few things we're still working on. 1. Cold starts are noticeably slower Right now, the slots start empty ( We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch. 2. Prompt processing is slow Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle. We're working on micro-batching prompt chunks ( 3. Speculative decoding gets weird with SSD offloading We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains. Right now, confidence-gated speculation ( Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me. Happy to answer questions or get into the (The second half of this post was written with some help from Claude.) [link] [comments] |
More from r/LocalLLaMA
-
Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
Oct 10
-
Fully local conversational AI: Whisper + Hermes 8B + Kokoro, zero cloud, running inside a plush toy
Oct 10
-
Typesafe ai raised 870m $ on hype (jev)
Oct 10
-
NVIDIA reportedly discontinuing RTX 5090, GB202 GPUs to be reserved for RTX PRO series
Oct 10
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.