§ tools · storyline
Adds Laguna on Apple GPUs via MLX engine
Ollama adds Laguna support on Apple GPUs via the MLX engine.
What's Changed Support Laguna on Apple GPUs via the MLX engine Quantize draft-model output heads at the requested type when creating speculative-decoding drafts. Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max).
Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4
§ sources1 publication · timeline below
- github.comOllama v0.32.4primary