DeepSeek publie V4-Flash-Vision-Exp, 305B multimodal en MIT
TL;DR
- DeepSeek ran a ten-day API-only window before releasing open weights, collecting production telemetry before committing the checkpoint publicly.
- Digital Applied's benchmark audit finds the model trails Opus 4.8 on 8 of 11 multimodal tests, directly undercutting the release's 'near-parity' framing.
- NYU Shanghai RITS notes some gains come from scoring a text-only baseline on multimodal suites; the 'Exp' designation signals DeepSeek's own production uncertainty.
DeepSeek a mis en ligne sur Hugging Face DeepSeek-V4-Flash-Vision-Exp, un modèle de 305 milliards de paramètres publié sous licence MIT. Sur ApexBench en Pass@1, il note 36,5, contre 26,2 pour DeepSeek-V4-Flash-0731 — et 39,4 pour Opus-4.8, la référence propriétaire retenue par la fiche modèle.
La fiche présente le travail sans détour : « We are excited to introduce DeepSeek-V4-Flash-Vision-Exp, our first experimental multimodal model in the DeepSeek-V4 family. » Elle précise que le modèle « builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities ». L'ossature réutilise les briques déjà nommées de la famille — DFlash attention, MoE, Hyper-Connections, DSpark forward path — sur lesquelles sont greffés les modules visuels.
Les gains portent d'abord sur les agents multimodaux. Agents' Last Exam passe à 27,3 (contre 25,2 pour la version précédente et 25,7 pour Opus-4.8). Chartography monte à 64,3, à sept dixièmes d'Opus-4.8 (65,0). Sur ZeroBench Pass@5, DeepSeek devance même la référence, 35,0 contre 34,0. Côté texte, DeepSeek revendique un profil « comparable » : Terminal Bench 2.1 à 83,9 (vs 82,7), DeepSWE 59,3 (vs 54,4), Toolathlon-Verified 75,9 (vs 70,3). L'écart reste net sur NL2Repo (57,7 contre 69,7) et DSBench-Hard (63,6 contre 71,7).
Le suffixe -Exp fait office d'avertissement : c'est le premier essai vision de la ligne V4. La fiche ne détaille ni la part active des paramètres du MoE, ni le corpus visuel du « continued training », ni la moindre évaluation de sécurité. C'est notre troisième publication DeepSeek en open-weights suivie ces dernières semaines, dans un contexte où Pékin ne réprime pas les open-weights de DeepSeek et Qwen.
Ce qu'en disent les autres médias
-
NYU Shanghai RITS Lire →
Technical deep-dive questioning the benchmark framing: gains partly reflect text-only baselines on multimodal suites; 'Exp' signals production uncertainty.
The vision variant maintains comparable performance on text-only agent tasks while improving on six of seven text benchmarks.
-
Digital Applied Lire →
Forensic pricing breakdown and independent benchmark audit showing the model trails Opus 4.8 on 8 of 11 tests extracted from DeepSeek's own announcement PNG.
Vision costs nothing extra per token. The rate card is identical to text-only deepseek-v4-flash.
-
Artificial Analysis Lire →
Independent ranking across 177 models: #5 in intelligence, #66 in speed at 108.7 tok/s, with output pricing above the median for its speed tier.
DeepSeek V4 Flash Vision (Reasoning, Max Effort) scores 51 on the Artificial Analysis Intelligence Index, placing it well above average.
-
OpenRouter Lire →
Provider availability snapshot: three inference providers already serve the model at $0.22 input/$0.66 output per million tokens, with 99.98-99.99% uptime.
Adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge.
-
Open Source For You Lire →
Open-source framing stressing MIT license scope, Hugging Face as deliberate distribution choice, and the 'Exp' label as an evolutionary rather than final release signal.
Shared on Bluesky by 1 AI expert
Article original publié par huggingface.co
Lire l'article original →Titre original : DeepSeek publie DeepSeek-V4-Flash-Vision-Exp sur Hugging Face, MoE multimodal 305 B pour agents visuels