r/LocalLLaMA · · 1 min read

Hot Expert Reload on GPU is what this community needs

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x 3090 speeds will be quite close to the full offload of these models to VRAM. This will make these almost SOTA models really usable locally.

submitted by /u/perelmanych
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA