huggingface.co web signal

AutoTrust Ships NVFP4 Build of GEV-26B-Decide at 17.1 GiB VRAM

TL;DR

  • The NVFP4 build cuts GPU memory for weights from 51.1 GiB to 17.1 GiB, a 66% drop, by quantizing only the routed-expert projections.
  • GPQA Diamond System 1 accuracy moves from 43.9% bf16 to 44.9% NVFP4, which the card calls within run-to-run noise of plus or minus 7 points over 198 questions.
  • The 26B MoE hits 95% success at about 85 ms per step on a 60-task computer-use benchmark using numbered boxes plus element text.

AutoTrust AI has posted an NVFP4-quantized build of its 26B decision model that cuts GPU memory for weights from 51.1 GiB to 17.1 GiB, a 66% drop, while leaving bf16 System 1 accuracy effectively unchanged on the two benchmarks the card publishes. On GPQA Diamond the System 1 number moves from 43.9% to 44.9% across 198 questions, which the card describes as within 'run-to-run noise (198 questions: ±7 points).' HLE text multiple choice shifts from 8.4% to 9.7% over 513 questions.

The quantization is targeted. Only the routed expert MLPs — 30 layers times 128 experts, 3,840 in total — move to NVFP4, since the card notes that 'routed experts hold ~22 GB of FP4-able bf16 out of 24 B parameters' while attention, dense MLP, and routers stay in bf16. Activations use static scales calibrated on 3,072 System 1 prompts.

On a single B200 running vLLM, System 1 posts a 45 ms 'Median latency: 45 ms for single request' and 257 decisions per second with 64 concurrent clients. On a 60-task computer-use benchmark described as 'numbered boxes + element text,' the model records 95% success at roughly 85 ms per step, versus about 260 ms for JEV-27B-VL at the same success rate. The card lists FlashInfer TRT-LLM NVFP4 MoE kernels for B200/B300/GB200, CUTLASS FP4 kernels for RTX PRO 6000 and RTX 5090, and a Marlin W4A16 fallback (weights FP4, activations bf16) for H100/H200/A100.

The card is explicit about what has not been retested: 'NVFP4: compared to bf16 only on GPQA Diamond and HLE; adaptive thinking not re-evaluated.' It also notes the robot-arm pick-and-place result — 40% success on 20 scenes at 61 ms per step — trails JEV-27B-VL at 75%. The release lands in a busy stretch for open-weights inference work, which our tracker has logged 93 times in the past 90 days.