v0.32.10
Mirrored from Ollama releases for archival readability. Support the source by reading on the original site.
What's Changed
- Models that don't set a
repeat_penaltynow default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself. - Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer.
- Fixed blob verification being skipped when an OCI manifest's config and layer share a digest.
New Contributors
- @vigneshakaviki made their first contribution in #15504
Full Changelog: v0.32.8...v0.32.10-rc1
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.