huggingface.co détecté sur le web

DeepSeek publie V4-Flash-Vision-Exp, 305B multimodal en MIT

6 médias qui couvrent ce sujet

TL;DR

  • The model uses 13B active parameters out of 284B total (sparse MoE) with a 1M-token context window, enabling fast inference without sacrificing capacity.
  • V4-Flash-Vision-Exp beats Opus 4.8 on Agents' Last Exam and ZeroBench Pass@5, but trails by 12 points on NL2Repo (57.7 vs 69.7), splitting 3 of 11 benchmarks.
  • Some of the multimodal benchmark advantage reflects V4-Flash's baseline ignoring image inputs entirely, inflating the apparent delta on vision tasks.

DeepSeek a mis en ligne sur Hugging Face DeepSeek-V4-Flash-Vision-Exp, un modèle de 305 milliards de paramètres publié sous licence MIT. Sur ApexBench en Pass@1, il note 36,5, contre 26,2 pour DeepSeek-V4-Flash-0731 — et 39,4 pour Opus-4.8, la référence propriétaire retenue par la fiche modèle.

La fiche présente le travail sans détour : « We are excited to introduce DeepSeek-V4-Flash-Vision-Exp, our first experimental multimodal model in the DeepSeek-V4 family. » Elle précise que le modèle « builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities ». L'ossature réutilise les briques déjà nommées de la famille — DFlash attention, MoE, Hyper-Connections, DSpark forward path — sur lesquelles sont greffés les modules visuels.

Les gains portent d'abord sur les agents multimodaux. Agents' Last Exam passe à 27,3 (contre 25,2 pour la version précédente et 25,7 pour Opus-4.8). Chartography monte à 64,3, à sept dixièmes d'Opus-4.8 (65,0). Sur ZeroBench Pass@5, DeepSeek devance même la référence, 35,0 contre 34,0. Côté texte, DeepSeek revendique un profil « comparable » : Terminal Bench 2.1 à 83,9 (vs 82,7), DeepSWE 59,3 (vs 54,4), Toolathlon-Verified 75,9 (vs 70,3). L'écart reste net sur NL2Repo (57,7 contre 69,7) et DSBench-Hard (63,6 contre 71,7).

Le suffixe -Exp fait office d'avertissement : c'est le premier essai vision de la ligne V4. La fiche ne détaille ni la part active des paramètres du MoE, ni le corpus visuel du « continued training », ni la moindre évaluation de sécurité. C'est notre troisième publication DeepSeek en open-weights suivie ces dernières semaines, dans un contexte où Pékin ne réprime pas les open-weights de DeepSeek et Qwen.

Ce qu'en disent les autres médias

Couverture consolidée 2h après publication

  1. DeepSeek API Docs Lire →

    First-party announcement with pricing mechanics: images tokenized at 384 tokens max billed at V4-Flash rates; new Files API enables image reuse via file_id.

    On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.
  2. The Next Web Lire →

    Critical benchmark audit: model wins only 3 of 11 head-to-heads vs Opus 4.8; flags that V4-Flash baseline ignoring images inflates the multimodal gain.

    This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge.
  3. OfficeChai Lire →

    Cost-as-strategy framing: the competitive frontier for Chinese labs is now cheap, fast vision models suited to production, not largest parameter count.

    Competition among Chinese AI labs isn't only about who has the biggest model, but about who can make a cheap, fast model good enough at seeing the world.
  4. Emergent Lire →

    Frames the experimental label as deliberate iteration strategy; highlights medical imaging, document processing, and quality control as target verticals.

    The experimental status allows for rapid iteration based on user feedback, while the Flash architecture suggests a focus on practical deployment scenarios.
  5. OpenRouter Lire →

    Infrastructure view: two providers live at launch with sub-2s latency, 99.94% uptime, and cache reads at $0.007/M tokens confirming day-one production routing.

Shared on Bluesky by 1 AI expert