shipfeedAI news, curated daily

09:30:32 CET
1 SEPT09:30:32shipfeed
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Fuses DFlash encoder into KV cache injection

DFlash optimizes encoder by fusing into KV cache injection to reduce device transfers.

yesterday · · primary fetch1 sourceupdated yesterday ·

spec : fuse the DFlash encoder into the KV cache injection (#27310) dflash : fuse the encoder into the KV injection decode The encoder is a single fc + norm, but running it as a separate llama_encode forced a device-to-host round trip of its output before the injection decode could re-upload it, plus a second graph build per round. Fold the encoder into the decoder's embd branch and feed the target features directly to one llama_decode. Assisted-by: Claude Fable nit Apply batched suggestions from code review Co-authored-by: Ruixiang Wang Fix missing references from renaming --------- Co-authored-by: Sigbjørn Skjæret Co-authored-by: Ruixiang Wang Website: Attestations: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.14) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Android: Android arm64 (CPU) Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows x64 (CUDA 12) - CUDA 12.4 DLLs Windows x64 (CUDA 13) - CUDA 13.3 DLLs Windows arm64…

read full article on github.com
§ sources1 publication · timeline below
  1. github.comllama.cpp b10715primary