A small sidecar model, released September 24, promises faster on-device vision AI

On September 24, 2026, Liquid AI released LFM2.5-VL-DSpark, an experimental "draft" model designed to make the company's LFM2.5-VL-3B vision-language model run substantially faster without changing what it outputs. A vision-language model, or VLM, is the kind of AI that can look at an image and answer questions about it: reading the total on a receipt, extracting numbers from a chart, describing a photo, or following a multi-turn conversation about a document. Those are exactly the workloads people might want on a personal device rather than in a data center, and exactly the ones where a three-billion-parameter model can feel slow.

The new release is a sidecar: a small model of about 280 million parameters, roughly 8.9% added on top of the 3B target model, that acts as a fast guesser during text generation. In Liquid AI's self-reported benchmarks, it made the target model's decoding phase up to 3.13x faster on a laptop-class Apple M5 Max and up to 2.66x faster on an NVIDIA H100 data-center GPU, with end-to-end gains reaching 2.62x and 2.27x respectively. Day-one support shipped in llama.cpp, MLX-VLM and SGLang, three widely used open-source inference tools, and the weights are available in both Safetensors and GGUF formats.

One caveat applies to every number in this article: they come from Liquid AI's own blog post and model card. No independent party has published its own measurements yet. The figures below are therefore labeled as vendor-reported throughout.

How speculative decoding works, in plain language

The technique behind DSpark is speculative decoding, and the core idea is easy to picture. Large language models generate text one token at a time, where a token is roughly a word or a chunk of a word. Each step requires a full pass through the big model, which is the slow part.

A draft model flips the workflow. Instead of asking the big model for every token, a small, cheap model first guesses a whole block of upcoming tokens, in this case blocks of eight or nine. Then the big model checks the whole block in a single pass, accepting the guesses that match what it would have produced and discarding the ones that do not. If the drafter is right often, the big model does far less work; if it is wrong, the corrections cost little because verification is cheap.

The crucial property is exactness. Liquid AI states the mechanism plainly in its blog post:

Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.

In other words, the speed comes from skipping redundant computation, not from cutting corners. Under greedy decoding the output is what the target model would have produced on its own. The model card adds that under matched sampling settings at non-zero temperatures, speculative decoding preserves the target model's output distribution. This is the mechanism as Liquid AI describes it; the claim that outputs are identical is well-grounded in how the method works, but as with the benchmarks, it is vendor-stated rather than independently verified.

What makes this release notable for vision models specifically is that the drafter must handle images as well as text. Liquid AI's solution is architectural: image patches and text tokens are projected into a shared representation before the layers the drafter taps, so the small model operates on hidden-state vectors of the same dimensionality no matter the input modality. That allowed the company to reuse the same recipe and inference algorithm it used for its text-only LFM2.5-DSpark drafters released in August 2026.

What the drafter actually is: 280M parameters and 8.9% extra memory

The drafter is deliberately tiny compared to its target. According to the model card, its 279.5 million parameters break down into a 193.0 million-parameter decoder stack of four full attention layers, a 21.0 million-parameter hidden-state projection, a 65.5 million-parameter Markov head, and small normalization and confidence components. It was trained for 10 epochs on vision-language supervised fine-tuning data weighted toward the workloads Liquid AI expects it to serve, and ablations across 3, 4 and 5 layers informed the final four-layer design.

The parameter budget matters for a practical reason: the drafter has to fit alongside the target model in memory. At 8.9% overhead on a 3B model, the cost is small enough that the combined system still runs on a laptop. The vocabulary is 128,000 tokens, and Liquid AI recommends a block size of 8 or 9 at inference depending on hardware, with Apple silicon using 8.

The numbers, all self-reported

All benchmarks follow the MMSpec benchmark methodology across six vision-based tasks: general VQA, text VQA, image captioning, chart VQA, complex reasoning and multi-turn conversation. All runs used batch size 1 at temperature 0, 16-bit processing, and measurements collected with Liquid AI's Pipette tool.

On-device, with MLX-VLM on an Apple M5 Max at block size 8, Liquid AI reports decoding that is 2.30x to 3.13x faster by task, with end-to-end latency improving by 1.56x to 2.62x. With llama.cpp on an Apple M3 Ultra, decoding improves by 1.57x to 2.14x and end-to-end by 1.30x to 1.77x.

On GPU, the model card's per-task table reports H100 decode speedups ranging from 2.04x (multi-turn) to 2.66x (COCO captioning), with end-to-end gains from 1.64x to 2.27x.

A data note on the H100 range: the blog post's summary sentence contains a likely typo, reading "20.4x to 2.66x faster decoding" for the H100 configuration. The model card's full benchmark table contradicts that figure, showing per-task decode speedups of 2.04x to 2.66x on the same hardware. This article reports the model card's consistent range of 2.04x to 2.66x and flags the blog's summary figure as an apparent typo.

The difference between decode and end-to-end numbers is not noise; it is the interesting part of the story, and Liquid AI is unusually candid about it.

The honest catch: Amdahl's law and why end-to-end gains are smaller

If decoding gets three times faster but half the total time is spent before decoding even starts, the whole job only gets somewhat faster. That is Amdahl's law: overall speedup is capped by the part of the workload that is not accelerated.

Liquid AI explains the problem directly:

Speculative decoding speeds up only decode, not vision encoding or prefill. When those stages already take up much of the wall time, even a large decode speedup gives only a modest end-to-end gain. This is Amdahl's law, where the overall speedup is capped by the part of the workload that isn't accelerated.

The company adds context on why this bites vision models especially hard: in ordinary language models, prefill is compute-bound and grows with prompt length, and VLMs add an image that first passes through a vision encoder and then enters the language backbone as hundreds of visual tokens alongside the text prompt. Edge devices have far less compute than data-center GPUs, so prefill and vision encoding take up a larger share of end-to-end latency there. The company notes one partial exception: the M5's per-core GPU neural accelerators narrow this gap on recent Apple silicon.

The practical takeaway for readers: the 3.13x headline is a decode-phase number on the most favorable hardware tested. If your workload involves very long prompts or many images relative to generated text, expect end-to-end gains closer to the lower bounds, such as 1.56x on the M5 Max or 1.64x on the H100.

Why this matters: private, responsive document and photo understanding

Running a vision-language model locally has two concrete benefits over calling a cloud API: latency and privacy. A model answering questions about a receipt, a chart, a contract page or a photo album on the user's own machine never uploads those images anywhere, and answers arrive without a network round trip. The barrier has been that capable open-weight VLMs are heavy and slow on consumer hardware.

A drafter that costs 8.9% extra memory but roughly halves to triples effective decoding speed on laptops directly addresses that barrier for the generation phase, which is typically the dominant cost in conversations and long-form outputs. Day-one support in llama.cpp, MLX-VLM and SGLang means developers do not need to write custom integration code; each project shipped support in its respective build, per pull requests referenced by Liquid AI (SGLang PR #40651, llama.cpp PR #29339, MLX-VLM PR #2280).

A few practical limitations are worth knowing. DSpark decoding in MLX-VLM currently uses greedy sampling, so users there must set temperature to 0. Each drafter checkpoint pairs with one specific target model. And the model itself is labeled experimental by its maker.

The bigger picture, and what remains uncertain

Liquid AI frames LFM2.5-VL-DSpark as an extension of a strategy it began with text models in August 2026: rather than shipping faster-but-lossier quantized models, it ships exact-speedup drafters that leave the target model untouched. The company describes its LFM2.5 family as open-weight and usable without restrictions, covering base, audio and vision variants on one architecture.

Whether the vendor-reported speedups hold up under independent testing, and whether similar drafters appear for other open model families, are the open questions worth watching. For now, the release is a concrete datapoint in a broader trend: the tools for running genuinely capable multimodal AI entirely on personal hardware are getting faster, cheaper in memory, and easier to install.

In my view, the most credible signal here is not the headline 3.13x figure but the pattern of the results: decode gains that are consistently large across tasks and hardware, paired with end-to-end gains the authors themselves caveat using Amdahl's law. Vendors rarely publish their own worst-case numbers so prominently, and the willingness to do so makes the rest of the benchmark table somewhat more believable, though still unverified.