Speeds up flash-attention on Apple Silicon with device tuning
Speeds up flash-attention on Apple Silicon with device tuning
metal : per-device tuned (Q, NE) for flash-attn vec (#26570) metal : per-device tuned (Q, NE) for flash-attn vec (#25750) rebase Q-generic FA vec body from 01dc93607 (#23114) add 53 f16 (Q,NE) flash-attn vec instantiations (vec 80 -> 133) add FA vec (Q,NE) tuning table + dispatch wiring + SMEM cap fallback add FA vec (Q,NE) perf sweep fill tuning result fold family table into a per-family representative SKU refactor tuning result format extend FA vec tuning to quantized KV caches sync fa vec tuner bucketing with runtime, use pointwise tuning regret update tuned table format and cleanup prefix fa_vec tuning procs with ggml_backend_metal_tuning_, drop unused fa_vec_override_active add device id -> token lookup for the offline tuning tool add ggml-metal-tuning skeleton add op-agnostic perf cell + median timing for the tuner add FA-vec graph build + tensor init to the tuner tools : add FA-vec (Q,NE) sweep, compression and table emit cool down and re-measure the dirty window on thermal drift test-backend-ops : replace the FA vec tune mode with a bounded (Q,NE) slice tools : document the Metal tuner, point the table comment at it abort on unknown KV type, single-source fa_vec_legal_ne…
- github.comllama.cpp b10615primary