Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the Measured on my M5, 24GB:
The same three cases on a 256 GB M3 Ultra hit 5.31–6.92 tok/s. Issues: long prompt prefill is trash (2,785 tokens is almost 9mins to first token), and it's text-only for now. Four model families now: Gemma 4 26B-A4B (~2 GB), Qwen 3.6 35B-A3B (~1.45 GB), DeepSeek-V4-Flash 284B-A13B (~6.8 GB), Inkling-Small 276B-A12B (~9.5 GB). I also got access to a few M3 Ultras, so I'll be testing and optimizing for higher configs too. But the primary goal stays the same: large MoE models on consumer grade hardware. Repo: https://github.com/NeelM0906/Mference — Swift + Metal, not a wrapper around MLX or llama.cpp. Mac app, CLI, and an OpenAI compatible server. Contributions welcome. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.