shipfeedAI news, curated daily

07:22:19 CET
8 SEPT07:22:19shipfeed
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Adds fused DeepSeek-V4 ops to Vulkan

Vulkan adds fused DeepSeek-V4 hyper-connection operations, previously available only in CUDA and Metal.

yesterday · · primary fetch1 sourceupdated yesterday ·

vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) (#26578) vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) CUDA has these ops from the DeepSeek-V4 merge and Metal gained them in PR 26459. Vulkan was the last major backend running the unfused primitive chain. On DeepSeek-V4-Flash the unfused Sinkhorn comb chain alone takes about 32% of decode op time on gfx1151 (Strix Halo), spread over roughly 16k dispatches per token. dsv4_hc_comb runs the full 20-iteration Sinkhorn in registers. A token's 4x4 comb matrix lives in 16 consecutive subgroup lanes, with idst in bits 0-1 and isrc in bits 2-3 to match the CPU reference layout, so subgroupShuffleXor by 1|2 reduces rows and by 4|8 reduces columns.

One dispatch replaces about 137 strictly ordered node executions per site. The shuffle masks never cross a 16-lane boundary, so a subgroup of size 64 packs 4 independent tokens. dsv4_hc_pre and dsv4_hc_post handle the elementwise stream collapse and fan-out, with per-token coefficients staged in shared memory. GGML_VK_DISABLE_DSV4_HC disables all three ops. The _COMB, _PRE and _POST variants gate each op independently so a single kernel can…

read full article on github.com
§ sources1 publication · timeline below
  1. github.comllama.cpp b10844primary