Comparing 4bit quants for MLX
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Curious what people think are the ideal 4-bit quantization types on MLX
These quants seem to be the most popular, at least for Gemma4 and Qwen3.6:
- OptiQ 4bit (mlx-community/Qwen3.6-27B-OptiQ-4bit)
- Unsloth dynamic 2.0 MLX (unsloth/Qwen3.6-27B-UD-MLX-4bit)
- oQ (Jundot/Qwen3.6-27B-oQ4)
- DWQ (can't find an example fo this one)
- native (mlx-community/Qwen3.6-27B-4bit)
Does anyone have any insight here?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.