v0.32.10
What's Changed - Models that don't set a `repeat_penalty` now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself. - Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer. - Fixed blob verification being skipped when an OCI manifest's config and layer share a digest. New Contributors * @vigneshakaviki made their first contribution in https://github.com/ollama/ollama/pull/15504 **Full Changelog**: https://github.com/ollama/ollama/compare/v0.32.8...v0.32.10-rc1
Why it mattersThis development may affect product decisions, technical choices or the direction of the AI market. Only one public source is currently available.