Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the Measured on my M5, 24GB:
The same three cases on a 256 GB M3 Ultra hit 5.31–6.92 tok/s. Issues: long prompt prefill is trash (2,785 tokens is almost 9mins to first token), and it's text-only for now. Four model families now: Gemma 4 26B-A4B (~2 GB), Qwen 3.6 35B-A3B (~1.45 GB), DeepSeek-V4-Flash 284B-A13B (~6.8 GB), Inkling-Small 276B-A12B (~9.5 GB). I also got access to a few M3 Ultras, so I'll be testing and optimizing for higher configs too. But the primary goal stays the same: large MoE models on consumer grade hardware. Repo: https://github.com/NeelM0906/Mference — Swift + Metal, not a wrapper around MLX or llama.cpp. Mac app, CLI, and an OpenAI compatible server. Contributions welcome. [link] [comments] |
More from r/LocalLLaMA
-
Hidden Reasoning from Claude and GPT are Decoded, and it is interesting
Aug 12
-
According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1
Aug 12
-
It's the final countdown, baby! Qwen is out in just over 7 hours!
Aug 12
-
FYI: Muse Glimmer Chat Template Got Updated Recently
Aug 12
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.