reuters.com web signal

DeepSeek V4-Flash is cheapest major model to run, firm finds

TL;DR

  • Artificial Analysis benchmarked DeepSeek V4-Flash at roughly 3 cents per test, more than 100 times cheaper to run than Anthropic's Claude Fable 5.
  • V4-Flash's headline API pricing is $0.14 per million input tokens and $0.28 per million output tokens.
  • V4-Flash scored 50 out of 100 on the Artificial Analysis Intelligence Index, matching Google's Gemini 3.6 Flash.

A benchmarking firm has handed the current cost-per-task lead to a Chinese open-weight model, and by a margin large enough to move budget conversations. According to Reuters reporting on figures from Artificial Analysis, DeepSeek's V4-Flash runs at roughly 3 cents per test, versus 86 cents for Kimi K3, $1.86 for OpenAI's GPT-5.6 Sol, and $3.15 for Anthropic's Claude Fable 5. Headline API pricing works out to $0.14 per million input tokens and $0.28 per million output tokens.

The interesting move in Artificial Analysis's methodology is that it isn't measuring sticker price alone, it's measuring what it costs to actually finish a task, which folds in how many tokens a model has to chew through to get there. As the firm frames it, a model with a low headline price can still prove expensive if it needs significantly more steps to produce an answer. On that combined measure V4-Flash lands more than 100 times cheaper than Claude Fable 5.

Capability-wise the picture is honest. V4-Flash scored 50 out of 100 on Artificial Analysis's Intelligence Index, matching Google's Gemini 3.6 Flash and finishing one point behind Meta's Muse Spark 1.1 and Z.AI's GLM-5.2. That places it squarely in the fast, cheap tier rather than at the frontier, so the like-for-like comparison is really against flash and mini tiers, not against the full GPT-5.6 or Claude Fable 5 that also appear in the price table.

The honest caveat is that the reporting treats a 'test' as a fixed unit without spelling out the underlying workloads, so a headline gap this dramatic is better read as a directional signal than a settled per-query cost for your app. What the reporting doesn't give you is latency, throughput under load, or how the model holds up on long-context work, and none of that shows in a cost benchmark.

If the numbers survive contact with real workloads, the practical shift is a new price floor for flash-tier calls. Teams sizing budgets around GPT-mini or Gemini Flash have a credible open-weight option, and resellers hosting V4-Flash get an obvious wedge against the incumbents.