huggingface.co web signal

Mistral posts NVFP4 build of Small 4 119B with vLLM, Red Hat

TL;DR

  • Mistral's Hugging Face page shows a new NVFP4-quantized checkpoint of Mistral Small 4 119B updated in the past day.
  • Small 4 is a 119B-total, 6.5B-active MoE with 128 experts, a 256K context, configurable reasoning and Apache 2.0 licensing.
  • The org's open-weights shelf now spans Ministral 3 8B, Medium 3.5 128B and Large 3 675B, several with NVFP4 siblings.

Mistral AI's Hugging Face organization page shows a fresh NVFP4-quantized build of Mistral Small 4 119B, updated in the past day, sitting at 1.48k downloads and 116 stars against the base model's 56.2k downloads and 421 stars.

The checkpoint is, per its own listing, "a post-training-activation quantized version of mistralai/Mistral-Small-4-119B-2603, created using llm-compressor as part of a collaboration with teams from vLLM & Red Hat", aimed at higher throughput and lower memory use. Small 4 itself was released on March 16, 2026 with "119B total parameters, 128 experts, 6.5B activated per token, a 256K context window, configurable reasoning, and an Apache 2.0 license".

The quantized upload joins a heavier shelf. The flagship Mistral-Medium-3.5-128B carries 124k downloads and 437 stars, described on the page as "Our first flagship models handling instruction-following, reasoning, and coding in a single set of opened-weights". Mistral-Large-3-675B-Instruct-2512, a granular MoE with "41B active parameters and 675B total parameters" that Mistral "trained it from scratch on 3,000 NVIDIA H200 GPUs", already has its own NVFP4 sibling with 9.81k downloads.

Neither the org page nor the search summaries publish accuracy numbers for the NVFP4 build against the base weights.

Shared on Bluesky by 1 AI expert