Speeds up AVX2 prompt processing for large batch IQ models
GGML releases AVX2 optimization for large batch prompt processing of IQ models.
AVX2: Speed up large batch size prompt processing of IQ models (#27402) Batched gemm for grid IQ quants Style updates and a bit more performance Clean up comments Move code around Vectorize IQ panel decode, lower threshold for speedup IQ panel: single-source gather layout, gate bias, vectorize interleave Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer Move IQ panel code out of repack into iqp.cpp, clean up comments Another comment sweep Add myself as iqp. codeownder Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size Renaming and moving The other half of renaming and moving Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition Update ggml/src/ggml-cpu/iqp.h Co-authored-by: Georgi Gerganov Add iqp_rows work buffer Revert "Add iqp_rows work buffer" This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.
Add NUMA fallback Add 10 row batch tests for IQP coverage on all grid IQ types Swap assert for return false in support check Move IQP mul_mat_id test --------- Co-authored-by: Georgi Gerganov Website: Attestations: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
- github.comllama.cpp b10726primary