§ feed · storyline
Liquid AI releases model using speculative decoding for 3.13x speed
Liquid AI releases LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its vision-language model, achieving up to 3.13× speedup on Apple silicon and 2.66× on H100 GPUs.
Liquid AI announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its vision-language model. MarkTechPost reports up to 3.13× decoding speedup on Apple silicon and 2.66× on an NVIDIA H100, with weights available on Hugging Face and support in SGLang, MLX-VLM, and llama.cpp.
§ sources1 publication · timeline below