Adds CUDA kernels for lightning indexer
Adds CUDA kernels for lightning indexer.
cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545) cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) chore : remove indentation of #pragma unroll cuda : remove unnecessary kernel template declarations cuda : add WARPS_PER_BLOCK and K_VECS_PER_BLOCK template parameters in lightning indexer kernels to avoid duplication of constants. cuda : relax MMA architecture requirements to Turing in lightning indexer implementation chore : renamed variables chore : rename ggml_cuda_op_lightning_indexer() to ggml_cuda_lightning_indexer() chore : TODO for AMD rocWMMA chore : whitespace formatting chore : another variable rename to fix problems caused by shadowing chore : yet another rename, this time uppercased all constants cuda : added alignment checks for Q and K tensors in lightning indexer implementation --------- Co-authored-by: Stanisław Szymczyk macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (ROCm 7.2)…
- github.comllama.cpp b10032primary