shipfeedAI news, curated daily

21:48:39 CET
13 AUG21:48:39shipfeed
pull to refreshlast sync
Just in — 8 new
§ topic

llama-cpp

29 stories · 7d·6 sources covering·30 active storylines

Updated Wed, 12 Aug 2026 CEST·29 new storylines this week·live

What this is

llama.cpp is an open-source C/C++ engine for running LLMs efficiently on local hardware. shipfeed tracks llama.cpp releases, new model support, and quantization and performance work.

storylines this week30 active

llama.cpp — Releases
LLAMA.CPP · b9496

Fixes Gemma 4 unified FPE on mtmd

mtmd: fix Gemma 4 unified FPE (#24088) macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu…

via github.com
Friday, April 3, 2026’s edition
Smol AI — Daily
AI · 1 source

not much happened today

Gemma 4 was launched by Google under an Apache 2.0 license, marking a significant open-model release focused on reasoning, agentic workflows, multimodality, and on-device use. It outperforms models 10x larger and has…

via news.smol.ai
Tuesday, June 2, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b9468

Adds real-time reasoning interruption via POST

server: real-time reasoning interruption via control endpoint (#23971) server: real-time reasoning interruption via control endpoint Builds on the manual reasoning budget trigger from #23949. Adds a CONTROL task that…

via github.com
Monday, June 1, 2026’s edition
Wednesday, May 13, 2026’s edition
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc17

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc27

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Saturday, August 8, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10328

Adds initial tool isolation support via Docker

server: add initial tool isolation support (via docker) (#26507) server: add initial tool isolation support (via docker) add docs adapt get_info py: fix type check cont separate tools_io_sandbox / tools_io_docker…

via github.com
Friday, August 7, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10322

Speeds up SYCL SSM_CONV by up to 1.85x

sycl: coalesce the ssm_conv window loads (#26612) test-backend-ops perf -o SSM_CONV on an Arc Pro B70, interleaved A/B against master, 6 reps, us/run: ne_a=[515,3328,1,1] ne_b=[4,3328,1,1] n_t=512 97.68 -> 52.95 1.85x…

via github.com
Tuesday, August 4, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10255

Extends SYCL oneDNN SDPA to non-FP16 KV caches

Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0–Q8_0 and FP32) (#25874) sycl: extend oneDNN SDPA to Q4_0-Q8_0 and F32 KV caches Extends the oneDNN SDPA path (PR #25222) to handle non-F16 KV caches by dequantizing…

via github.com
Sunday, August 2, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10227

Adds Qwen3 specialized chat parser

chat : add qwen3 specialized parser (#26252) Add tagged thinking tool parser chat : refactor and add permute helper cont : add support for omission cont : update tool delimiters cont : add comment for qwen3-coder cont…

via github.com
Wednesday, July 22, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10087

Adds support for Laguna XS.2 & M.1

Add support for Laguna XS.2 & M.1 (#25165) Website: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64…

via github.com
Friday, July 3, 2026’s edition
Monday, June 29, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b9840

Adds DeepSeek V4 model support with graph optimization

DeepSeek V4 (#24162) convert: add dsv4 conversion add basic setup add llm_graph_input_dsv4 add save-load state add sinkhorn eps - correction by @fairydreaming add rope fix cleanup dead code fix bugs support pro model…

via github.com
Wednesday, June 3, 2026’s edition
Saturday, May 16, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · 1 source

llama.cpp b9186

sync : ggml macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan)…

via github.com
Wednesday, May 13, 2026’s edition
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc21

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc31

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc20

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc29

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc15

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Ollama — Releases
OLLAMA · 1 source

Ollama v0.30.0-rc22

This version of Ollama will change the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX is used to accelerate model inference on…

via github.com
Yesterday’s edition
llama.cpp — Releases
LLAMA.CPP · b10369

Adds pocket-tts with 80% speedup on CUDA, 50% on CPU

mtmd: support pocket-tts (#26871) adapt the api text model ok working impl, need verify and clean up mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so…

via github.com
Tuesday, August 11, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10355

Adds multi-output backend sampling with token speculation

llama : support multi-output backend sampling (#25532) Enable backend sampling with token speculation Clamp the mask sum before converting it into the sampled index Add a numeric context parameter declaring the maximum…

via github.com
Monday, August 10, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10336

Refactors wgsl and simplifies flash_attn in ggml-webgpu

ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134) Website: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…

via github.com
Sunday, August 9, 2026’s edition
llama.cpp — Releases
LLAMA.CPP · b10331

Fixes get_info to report the isolate working directory

server: report the isolate working directory from get_info (#26773) server: report the isolate working directory from get_info Without an explicit cwd, get_info fell back to the server process working directory even…

via github.com