shipfeedAI news, curated daily

06:38:24 CET
25 AUG06:38:24shipfeed
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Speeds Metal op compilation with per-op source split

GGML speeds Metal kernel compilation with per-op source splitting and parallel builds.

yesterday · · primary fetch1 sourceupdated yesterday ·

metal: per-op source split + parallel compile (#26561) metal : per-op source split + parallel compile (#24021) preliminary extract common header op source split split metallib into 8 libs && load in parallel derive kernel->library routing from functionNames x-macro lib list + underscore filenames, dedup QK_NL, MRC fixes op source split 8 to 20 improve robustness of source fallback clean up change bool -> atomic_bool only prepend headers that source actually includes no semaphore, use GCD global queue dedup library compile path, fix NSError lifetime, rename gla relocate upstream concat/rope_back/repeat kernel changes into split files move ggml-common.h from common.h into dequantize.h to shrink binary size --------- Co-authored-by: lvyichen metal: add col2im_1d op (f32/f16/bf16) (#25176) metal : add set_rows with src0 f16 (#25434) metal : add CONV_2D_DW (depthwise convolution) support (#21565) metal : add Q2_0 support (#25419) metal: fuse snake activation (mul, sin, sqr, mul, add) (#25459) ggml-metal: FWHT kernel for metal backend (#25924) metal : port new kernels into the split sources Move the kernels added on master after the split (lightning indexer, DSv4 hyper-connections…

read full article on github.com
§ sources1 publication · timeline below
  1. github.comllama.cpp b10614primary