r/MachineLearning · · 1 min read

Proposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

I ported MoE expert expansion to llama.cpp 🚀

Run MoE models with MORE routed experts than the native top-K (8->x), adaptive threshold, 99→50% influence decay, layer range. Runtime-only, all backends.

Tested on Qwen 3.6 35B A4B+

https://github.com/vagrillo/llama.cpp/blob/moe-expansion/docs/moe-expansion.md

submitted by /u/Specific-Tax-6700
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning