82 TPS On Qwen 3.6 27b On A Macbook Pro | Introducing MTPLX V2: The Fastest Way To Run MLX Models.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey Everyone, here is an update on MTPLX! One month after releasing MTPLX V1 which brought a swift based app and upgraded CLI for coding use I am happy to announce MTPLX V2. The biggest change is Turbo Mode: using custom verify-specialized quantized-matmul kernels plus a compiled verify step we have achieved 82 TPS on a Macbook pro m5 max at a temperature of 0.6 We also released significant changes to SSD KV cache and long context tool calling improvements. here are the preliminary benchmarks from Ivan Fioravanti showing MTPLX vs oMLX vs DGX spark. Looking forward to hearing everyone’s thoughts on the fastest MLX runtime. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.