huggingface.co détecté sur le web

DeepSeek publie V4-Flash-Vision-Exp, 305B multimodal en MIT

6 médias qui couvrent ce sujet

TL;DR

  • DeepSeek ran a ten-day API-only window before releasing open weights, collecting production telemetry before committing the checkpoint publicly.
  • Digital Applied's benchmark audit finds the model trails Opus 4.8 on 8 of 11 multimodal tests, directly undercutting the release's 'near-parity' framing.
  • NYU Shanghai RITS notes some gains come from scoring a text-only baseline on multimodal suites; the 'Exp' designation signals DeepSeek's own production uncertainty.

DeepSeek a mis en ligne sur Hugging Face DeepSeek-V4-Flash-Vision-Exp, un modèle de 305 milliards de paramètres publié sous licence MIT. Sur ApexBench en Pass@1, il note 36,5, contre 26,2 pour DeepSeek-V4-Flash-0731 — et 39,4 pour Opus-4.8, la référence propriétaire retenue par la fiche modèle.

La fiche présente le travail sans détour : « We are excited to introduce DeepSeek-V4-Flash-Vision-Exp, our first experimental multimodal model in the DeepSeek-V4 family. » Elle précise que le modèle « builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities ». L'ossature réutilise les briques déjà nommées de la famille — DFlash attention, MoE, Hyper-Connections, DSpark forward path — sur lesquelles sont greffés les modules visuels.

Les gains portent d'abord sur les agents multimodaux. Agents' Last Exam passe à 27,3 (contre 25,2 pour la version précédente et 25,7 pour Opus-4.8). Chartography monte à 64,3, à sept dixièmes d'Opus-4.8 (65,0). Sur ZeroBench Pass@5, DeepSeek devance même la référence, 35,0 contre 34,0. Côté texte, DeepSeek revendique un profil « comparable » : Terminal Bench 2.1 à 83,9 (vs 82,7), DeepSWE 59,3 (vs 54,4), Toolathlon-Verified 75,9 (vs 70,3). L'écart reste net sur NL2Repo (57,7 contre 69,7) et DSBench-Hard (63,6 contre 71,7).

Le suffixe -Exp fait office d'avertissement : c'est le premier essai vision de la ligne V4. La fiche ne détaille ni la part active des paramètres du MoE, ni le corpus visuel du « continued training », ni la moindre évaluation de sécurité. C'est notre troisième publication DeepSeek en open-weights suivie ces dernières semaines, dans un contexte où Pékin ne réprime pas les open-weights de DeepSeek et Qwen.

Ce qu'en disent les autres médias

Couverture consolidée 24h après publication

  1. NYU Shanghai RITS Lire →

    Technical deep-dive questioning the benchmark framing: gains partly reflect text-only baselines on multimodal suites; 'Exp' signals production uncertainty.

    The vision variant maintains comparable performance on text-only agent tasks while improving on six of seven text benchmarks.
  2. Digital Applied Lire →

    Forensic pricing breakdown and independent benchmark audit showing the model trails Opus 4.8 on 8 of 11 tests extracted from DeepSeek's own announcement PNG.

    Vision costs nothing extra per token. The rate card is identical to text-only deepseek-v4-flash.
  3. Artificial Analysis Lire →

    Independent ranking across 177 models: #5 in intelligence, #66 in speed at 108.7 tok/s, with output pricing above the median for its speed tier.

    DeepSeek V4 Flash Vision (Reasoning, Max Effort) scores 51 on the Artificial Analysis Intelligence Index, placing it well above average.
  4. OpenRouter Lire →

    Provider availability snapshot: three inference providers already serve the model at $0.22 input/$0.66 output per million tokens, with 99.98-99.99% uptime.

    Adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge.
  5. Open Source For You Lire →

    Open-source framing stressing MIT license scope, Hugging Face as deliberate distribution choice, and the 'Exp' label as an evolutionary rather than final release signal.

Shared on Bluesky by 1 AI expert