AI models / The AI briefing

Liquid AI releases an experimental drafter for its vision-language model

A smaller model drafts tokens and a larger one checks them. Liquid AI is testing this route to faster vision-language inference.

Comparison of two model responses with an example image.
From Hugging Face. Static frame from the source demonstration.

Liquid AI released LFM2.5-VL-DSpark, an experimental draft model for LFM2.5-VL-3B. It uses speculative decoding to propose tokens that the main model checks, adding memory overhead in exchange for faster decoding in the developer’s tests. The release supports llama.cpp, MLX-VLM and SGLang.

How the experimental drafter works

Liquid AI describes DSpark as a draft model for LFM2.5-VL-3B. It proposes blocks of tokens that the main model verifies, using shared representations for image and text input. This is an inference technique intended to reduce the work needed to generate a response.

The project discusses integration with inference runtimes including llama.cpp, MLX-VLM and SGLang. Adding a drafter also introduces memory overhead, so performance must be considered alongside the resources needed to run both components.

Why the workload matters

Faster token decoding does not necessarily mean the entire request is faster by the same amount. Image encoding and processing the input can dominate some workloads. The reported measurements are developer tests of an experimental system, not a guarantee for every device or application.

Original source

This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.

Read the original at Hugging Face

← Back to all AI news