you can now use MTP in GLM-Air
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| If anyone still remembers GLM-4.5-Air from last year, you can now get a nice speedup by enabling MTP in llama.cpp. It is a 106B MoE with only 12B active parameters, which makes it interesting for machines with lots of memory but limited compute, such as Strix Halo or DGX Spark. I use it on 3090s. It's still great for creative writing, especially since we never got Gemma 4 124B MoE. There are multiple creative-writing / RP finetunes available on Hugging Face: https://huggingface.co/models?other=base_model:finetune:zai-org%2FGLM-4.5-Air&sort=likes (some even from this year). I also recommend Intellect 3.x by PrimeIntellect If your GGUF does not include the MTP block, you can download a small file from here: https://huggingface.co/jacek2024/GLM-4.5-Air-MTP-GGUF Thanks a lot to devMiikaK and HeadCutter for testing the PR while it was in progress. PS. It also works for the full GLM-4.5, but I doubt anyone still uses it ;) [link] [comments] |
More from r/LocalLLaMA
-
Any upcoming models to be excited about?
Aug 23
-
Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system
Aug 23
-
We quantized Qwen 3.8 27B and compared the quants on an RTX 6000
Aug 23
-
Nvidia Customers Notified About AI-Related Price Hikes Above 15%
Aug 23
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.