Ollama releases · · 1 min read

v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)

Mirrored from Ollama releases for archival readability. Support the source by reading on the original site.

  • mlx: speed up Qwen 3.8 prompt processing

Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU.

  • address comments

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Ollama releases