liquid.ai web signal

Liquid AI Opens d1-3B and d1-omni-600M Decision Models

TL;DR

  • d1-3B scores 48.57 on Decision Index v0.2.1 and lands on par with Decider 35B-A3B, a model 12x its size.
  • d1 models return an answer in a single forward pass with no tokens, running in 8 ms on an NVIDIA RTX 4090.
  • Both weights ship on Hugging Face with native llama.cpp support across Apple, AMD, Qualcomm, and NVIDIA platforms.

Liquid AI's new d1-3B scores 48.57 on the Decision Index v0.2.1, a benchmark the company controls, which it says puts the model "ahead of every model under 10B and on par with Decider 35B-A3B, a decision model 12x its size." Its smaller sibling d1-omni-600M scores 15.95 on the same split.

The models do not generate text. "Unlike our generative Liquid Foundation Models (LFMs), our d1 decision models don't produce tokens. Instead, they produce an answer in a single forward pass," Liquid AI writes in the release post. d1-3B inherits from LFM2.5-VL-3B and takes text plus images. d1-omni-600M is built on LFM2.5-Encoder-350M with added vision and audio encoders.

The pitch is latency. A single question runs in 8 ms on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano. A 384-pixel image lifts those to 17 ms and 202 ms at the two extremes. The use cases Liquid AI floats are "gesture-controlled games" and "live content moderation" running frame-by-frame on camera input.

Both models are on Hugging Face with native llama.cpp support across Apple, AMD, Qualcomm, and NVIDIA platforms, described as open-weight and free to "download, fine-tune, and deploy without restrictions." The post does not name the specific license text that governs that promise. It lands in a busy week for small-model releases: Perplexity shipped its pplx-embed-v2-late embeddings the same day, and llama.cpp itself picked up GLM-5.3-Flash last week.