shipfeedAI news, curated daily

03:10:38 CET
28 SEPT03:10:38shipfeed⋯
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Speeds up SYCL FWHT for block widths 1024-8192

ggml adds SYCL FWHT kernels for block widths 1024-8192, replacing dense GEMM fallback.

yesterday · · primary fetch1 sourceupdated yesterday ·

sycl: FWHT kernels for block widths above 512 (#29243) The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n). fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers.

Same butterfly and sign convention as the existing narrow kernel. ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a GPU-labeled platform to exist and throws before any op-level test can run, so test-backend-ops could not be exercised on this box (a GPU-less pod) even via the CPU device. Verified instead with a standalone harness: the same kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1, OpenCL CPU backend), checked against an independent…

read full article on github.com ↗
§ sources1 publication · timeline below
  1. github.comllama.cpp b11216primary