AI that decides, not writes
Liquid AI, the startup behind the Liquid Foundation Models line, released two open-weight 'decision' models on Hugging Face on October 7, 2026: d1-3B and d1-omni-600M. The release matters for what the models do not do. They do not generate text. You hand one a 'state', which can be a piece of text, a JSON record, a photo, or in the smaller model's case a short audio clip, together with a set of named questions, and it returns structured answers in a single forward pass through the network. No tokens are produced.
That distinction is worth unpacking for non-specialists, because most of the AI systems people hear about, including chatbots, work by generating tokens. A token is a small chunk of text, and the model produces them one at a time, each step requiring another pass through the network. That is why a chatbot answer takes seconds: the model is effectively writing word by word, and each written word costs compute. A decision model skips that loop entirely. The answer, whether a yes/no, a pick from a list of named options, or a score on a defined scale, is read directly from the model's internal probability distribution over the options. There is nothing to write, nothing to parse, and no chance of the model rambling.
The practical consequence is latency and cost. Liquid AI reports that d1-3B answers a question in 16 milliseconds on an NVIDIA Jetson AGX Thor and 50 milliseconds on a Jetson Orin Nano, the kind of small boards hobbyists and industrial builders actually buy. On desktop GPUs, the company reports 8 milliseconds on an NVIDIA RTX 4090 and 9 milliseconds on an AMD MI325X, and 30 milliseconds on an Apple M5 Pro laptop. These are vendor-reported numbers, measured in collaboration with NVIDIA, and no independent party has verified them yet.
The use cases Liquid AI names are the unglamorous plumbing of software: routing support tickets to the right team, scoring how urgent a request is, detecting toxic comments, classifying user intent, checking whether an extraction is correct, judging and reranking outputs from other models, guarding agents, and visual inspection. The company's example code shows exactly this pattern: one customer message, three questions asked of it at once (is the customer asking for a refund, which team should handle it, how urgent is it), answered in a single pass. The workloads are decided, not written.
Two models, two very different backbones
Both models are built on Liquid AI's own foundation models, but from two very different starting points. d1-3B is post-trained from LFM2.5-VL-3B, the company's 3.1-billion-parameter vision-language model. d1-omni-600M is post-trained from LFM2.5-Encoder-350M, a bidirectional encoder, with a 94-million-parameter vision encoder and a 112-million-parameter audio encoder added on top. Its total parameter count is 587 million, which the model card describes as a shared trunk plus the two modality encoders.
The size difference maps to a modality difference. d1-3B handles text and images in the same state. d1-omni-600M handles either text plus images, or text plus up to 30 seconds of audio, not both at once; the model card states that a request carrying both images and audio raises an error. The audio capability, per the model card, was trained on recordings of an English speaker interacting with an assistant, with tasks covering what kind of utterance it is, its topic, and what the speaker wants. That is a narrow audio domain, and the model card flags it plainly.
Liquid AI describes d1-omni-600M as an early research release that is still under active development. That caveat is not cosmetic. The company reports no latency numbers for the smaller model at all in this release, and no standalone vision or audio benchmarks for either model. The stated reason: the Decision Index v0.3 includes only a private vision split, and audio decision benchmarks are, in the company's words, currently an open problem. So anyone choosing between the two models today is comparing on text benchmarks plus vibes on the multimodal side.
What the benchmarks say, and who ran them
The headline benchmark claim is vendor-run and should be read as such: Liquid AI says d1-3B scores 48.57 on the Decision Index 0.2.1, which the company positions as the best score of any decision model under 10 billion parameters, ahead of Decider 35B-A3B at 47.11, a model more than ten times its size. The Decision Index is a public leaderboard for models of this type, but Liquid AI scored its own models with the official scorer rather than submitting to the leaderboard, and other rows come from the public leaderboard v0.2.1. Independent verification is absent.
Beyond the index, the release blog reports results on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical question answering, and cross-lingual understanding. Per the blog post, d1-3B achieves a mean score of 82.9, above Decider 4B at 81.1 and Decider 2B at 77.1, with individual dataset scores of 83.3 on SQuAD 2.0, 93.3 on Civil Comments, 86.9 on MASSIVE intent, 68.3 on PubMedQA, 86.3 on BoolQ, 85.6 on XNLI, and 76.4 on PAWS-X. The company reports d1-omni-600M at a mean of 78.4, which it says surpasses Decider 2B at 77.1 with roughly a quarter of the parameters. Its strongest single result is 95.8 on Civil Comments toxicity detection, and its weakest is 61.3 on PubMedQA.
One consistency note that a careful reader should know: the d1-omni-600M model card's table shows slightly different per-dataset scores for d1-3B than the blog post does (for example 85.3 rather than 83.3 on SQuAD 2.0), while the mean of 82.9 matches in both. The blog post is the primary release document cited here, and the discrepancy does not change any headline claim, but it is the kind of detail that independent benchmarks eventually settle. There is also a genuine tension in the messaging: the release blog says d1-omni-600M scores 78.4, surpassing Decider 2B at 77.1, while the model card's own Decision Index table lists d1-omni-600M at only 15.95 on that index, far below every other listed model. These measure different things: the 78.4 is the mean across the seven public datasets, and the 15.95 is the broader multi-category Decision Index, where the tiny model trails badly on knowledge, tools and arts. Both numbers come from the vendor, and both should be understood before anyone picks a model.
The vision side has one positive vendor-reported datapoint: the model card states that d1-3B scores 74.1 on 11 public image benchmarks, marginally above its LFM2.5-VL-3B backbone at 73.9, indicating that post-training for decisions did not destroy the base model's vision ability. Again, this is Liquid AI's own evaluation.
Why it matters, and the fine print
For a builder, the pitch is concrete. A small business could run ticket routing and urgency scoring on a laptop or a few-hundred-dollar Jetson board with no cloud round trip, no per-token bill, and no data leaving the premises. A moderation pipeline could classify comments in tens of milliseconds. A voice application could route spoken commands on device. The weights are downloadable on Hugging Face, the models run through Hugging Face transformers (the model cards specify version 5.14 or newer for d1-3B and 5.15 or newer for d1-omni-600M), and the company points to a System One Arcade Space with ten demos.
The costs and caveats are equally concrete. These models require trusting code shipped with the weights (trust_remote_code=True), which is normal for this family but is a supply-chain consideration. The licenses are listed as 'other' rather than a standard open license, so commercial users should read the LICENSE file before shipping. Latency figures exist only for d1-3B. And every quality number in this release comes from the vendor that made the models.
For non-specialists, the deeper point is architectural. The industry has spent three years making token generation faster and cheaper, and that progress is real. But a large class of everyday AI work was never about writing at all: it is about deciding. Turning that class of work into a single forward pass with a fixed output structure is a different bet, and one that trades generality (these models cannot chat) for speed, determinism and small footprint. Whether decision models become a standard pipeline component or a niche depends on whether the quality numbers hold up under independent testing, and on how quickly the open benchmark ecosystem for multimodal decisions matures, which Liquid AI itself says is an open problem.
