Stabilizes sampling graph topology
Llama stabilizes sampling graph topology by ensuring all samplers build consistent output chains.
llama: keep the backend sampling graph static across ubatches (#30223) The reserve builds n_outputs_max_per_seq sampling chains per sampler, while a decode built one per output row, so the graph changed its topology after the reserve and GGML_SCHED_NO_REALLOC builds aborted on the next same sized graph. Every sampler now builds n_outputs_max_per_seq chains, the ones without a row of the ubatch on the padding row and not selected, and graph_max_nodes counts them. Website: Attestations: macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries Ubuntu x64 (ROCm 10.0) Ubuntu x64 (OpenVINO) Ubuntu x64 (SYCL FP32) Ubuntu x64 (SYCL FP16) Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Android: Android arm64 (CPU) Android arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide Windows: Windows x64 (CPU) Windows arm64 (CPU) Windows arm64 (OpenCL Adreno) Windows…
- github.comllama.cpp b11530primary