shipfeedAI news, curated daily

14:21:02 CET
29 JUL14:21:02shipfeed
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Speeds up Mamba-2 prefill via chunked SSD matmul

ggml adds chunked SSD matmul optimization for Mamba-2 model prefill acceleration.

yesterday · · primary fetch1 sourceupdated yesterday ·

ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675) ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC. ggml-cuda: review comments fixed. ggml-cuda: Fuse M matrix materialization into pre_matmul kernel and enabled test. ggml-cuda: test updates and fixes ggml-cuda: test updates to remove hardcoding of tensor initialise data limits. ggml-cuda: ssd minor review comment fixed. ggml-cuda: ssd minor CICD fixed. CUDA SSD: Fixes correctness by promoting s0_stride_seq to int64_t, improves memory coalescing in ssm_ssd_prepare_dt_kernel, and boosts efficiency by merging B_weighted and C_scaled; also addresses prior review comments.

cuda: fix sdata read-write race in prepare_dt fallback scan loop Website: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.2) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Android: Android arm64 (CPU) Windows: Windows x64 (CPU) Windows arm64 (CPU)…

read full article on github.com
§ sources1 publication · timeline below
  1. github.comllama.cpp b10164primary