r/LocalLLaMA · · 1 min read

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini

I went with the recommended settings for the best generation quality and ran a few benchmarks:

  • temperature: 0.7
  • top_p: 0.95
  • top_k: 40
  • reasoning_effort: high

I must say, the outcome is not bad at all - really good generation speed and prompt processing, okay memory footprint and good quality across the board. Will for sure give it a try to fuel my agents and might also try to do some coding with it.
All benchmarks run I did you can find here: https://llm-bench.io/models/nex-n2-5-mini-mlx-4bit

Quant I used: https://huggingface.co/abenzerps/Nex-N2.5-mini-MLX-4bit

submitted by /u/DerTomsn
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA