Speeds up Qwen3.5 on Apple GPUs with speculative decoding
Ollama updates MLX engine for faster Qwen3.5 on Apple GPUs using speculative decoding and improves OpenAI compatibility.
What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically `/v1/chat/completions` streaming now matches OpenAI's wire format: `role` only on the first chunk, `finish_reason` on its own chunk, and usage in a separate chunk with `stream_options.include_usage`. Truncated OpenAI responses now report `finish_reason: "length"` instead of `"tool_calls"`. `ollama run kimi-k3` now offers `kimi-k3:cloud` for cloud-only models that publish no default tag, instead of failing.
TUI fixes: pipe-delimited prose no longer renders as a table, Enter accepts the highlighted `@` file completion, and `/prompt` scrolling is no longer laggy. Experimental image generation has been temporarily removed. Continue using 0.32.5 for image generation support Updated the MLX and llama.cpp engines. Full Changelog: https://github.com/ollama/ollama/compare/v0.32.5...v0.32.6-rc0
- github.comOllama v0.32.6-rc0primary