
Liquid AI released LFM2.5-VL-DSpark, an experimental draft model for LFM2.5-VL-3B. It uses speculative decoding to propose tokens that the main model checks, adding memory overhead in exchange for faster decoding in the developer’s tests. The release supports llama.cpp, MLX-VLM and SGLang.
How the experimental drafter works
Liquid AI describes DSpark as a draft model for LFM2.5-VL-3B. It proposes blocks of tokens that the main model verifies, using shared representations for image and text input. This is an inference technique intended to reduce the work needed to generate a response.
The project discusses integration with inference runtimes including llama.cpp, MLX-VLM and SGLang. Adding a drafter also introduces memory overhead, so performance must be considered alongside the resources needed to run both components.
Why the workload matters
Faster token decoding does not necessarily mean the entire request is faster by the same amount. Image encoding and processing the input can dominate some workloads. The reported measurements are developer tests of an experimental system, not a guarantee for every device or application.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at Hugging Face