shipfeedAI news, curated daily

18:18:26 CET
25 AUG18:18:26shipfeed
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Aligns /v1/chat/completions streaming with OpenAI's wire format

Ollama updates /v1/chat/completions streaming to match OpenAI's wire format.

Aug 4 · · primary fetch1 sourceupdated Aug 4 ·

What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically `/v1/chat/completions` streaming now matches OpenAI's wire format: `role` only on the first chunk, `finish_reason` on its own chunk, and usage in a separate chunk with `stream_options.include_usage`. Truncated OpenAI responses now report `finish_reason: "length"` instead of `"tool_calls"`. `ollama run kimi-k3` now offers `kimi-k3:cloud` for cloud-only models that publish no default tag, instead of failing.

TUI fixes: pipe-delimited prose no longer renders as a table, Enter accepts the highlighted `@` file completion, and `/prompt` scrolling is no longer laggy. Experimental image generation has been temporarily removed. Continue using 0.32.5 for image generation support Updated the MLX and llama.cpp engines. Full Changelog: https://github.com/ollama/ollama/compare/v0.32.5...v0.32.6-rc0

read full article on github.com
§ sources1 publication · timeline below
  1. github.comOllama v0.32.6primary