Running Qwen3.8-Flash-Next locally on a 12GB VRAM card
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Now that the dust has settled a bit - here's a write-up on running Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on relatively middle-tier hardware (RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe on Linux). I started out with bare 6 tok/s and through latest patches and optimizations getting close to 20 tok/s generation. You just need enough RAM. For me this is the most intelligence possible on this machine right now. The 27B dense is not a choice because of low VRAM but may make more sense for other configs like 24GB VRAM owners. It actually surpasses the 27B model on most tasks as well so its great for Low VRAM, High/fast RAM configs. PP is still a bit low at 300-350 tok/s. What helped - Master branch (19.35 t/s): Latest commit with MoE improvements. - MTP Variant - PR #28243 + Compact MTP (20.65 t/s): MTP support is not yet merged so need to apply this PR enables Daniel Han's 1.78 GB `shared-Q4_K_M` compact head. Combined with `-ncmoe 45`, it yields 77–96% acceptance and breaks through the 20 t/s barrier on every tested task (coding, summarization, creative). With such low VRAM, MTP is not a huge jump because you have to give up a few layers to store the MTP head in VRAM. Only the shared + Q4_K_M in MTP gets a beneficial uptick. Using commercial models to research, optimize and benchmark inference for local models helps a ton (GLM 5.3 flash with opencode go, so did Astra, Gemini 3.8 etc.) Lot more details in the post (AI-assisted). [link] [comments] |
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.