shipfeedAI news, curated daily

03:11:14 CET
28 SEPT03:11:14shipfeed⋯
pull to refreshlast sync
Just in — 30 new
§ local-llm · storyline

Enables long-document and multimodal reranking with causal LLMs

Llama.cpp enables long-document and multimodal reranking with causal LLMs.

today · · primary fetch1 sourceupdated today ·

server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876) server : allow splitting RANK pooling for causal LLM rerankers Rerank models fall into two categories: bidirectional cross-encoders (BERT, etc.) that require all tokens in a single physical batch, and causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use chunked prefill like any other decoder. Previously the server rejected all RANK-pooling inputs larger than n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to determine last-token pooling. This broke long-document and multimodal reranking for causal models.

Fix: expose llama_get_causal_attn(ctx) so the server can check the effective runtime attention type (reflecting any --attention override or set_causal_attn call). Also expose llama_model_is_causal(model) for querying the static architectural property from GGUF metadata. can_split() now permits chunked prefill for RANK pooling when the context is causal. The graph builder's inline arch check is replaced with the same cparams.causal_attn predicate, removing the duplication. Assisted-by: Opencode/Qwen3.8-27B remove unused llama_model_is_causal…

read full article on github.com ↗
§ sources1 publication · timeline below
  1. github.comllama.cpp b11223primary