Enables long-document and multimodal reranking with causal LLMs
Llama.cpp enables long-document and multimodal reranking with causal LLMs.
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876) server : allow splitting RANK pooling for causal LLM rerankers Rerank models fall into two categories: bidirectional cross-encoders (BERT, etc.) that require all tokens in a single physical batch, and causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use chunked prefill like any other decoder. Previously the server rejected all RANK-pooling inputs larger than n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to determine last-token pooling. This broke long-document and multimodal reranking for causal models.
Fix: expose llama_get_causal_attn(ctx) so the server can check the effective runtime attention type (reflecting any --attention override or set_causal_attn call). Also expose llama_model_is_causal(model) for querying the static architectural property from GGUF metadata. can_split() now permits chunked prefill for RANK pooling when the context is causal. The graph builder's inline arch check is replaced with the same cparams.causal_attn predicate, removing the duplication. Assisted-by: Opencode/Qwen3.8-27B remove unused llama_model_is_causal…
- github.comllama.cpp b11223primary