Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: Ideal 4-bit quant ≈ 82 GB The big n-gram table is sparsely accessed → excellent candidate for system RAM offload. This architecture could be surprisingly local-friendly once the weights drop. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.