deepmind.google web signal

Google Ships Gemini 3.7 Flash, Coding Bench Jumps to 43.6%

4 sources tracking this story
Google Coding Tools ai-business

TL;DR

  • The 50% intro price expires January 1, 2027, doubling per-token cost for any team that scales production usage on it now.
  • DeepSWE v1.1 jumped from 49.0% to 65.3% over the prior Flash, per vendor-reported benchmarks against Claude Sonnet 5 and GPT-5.6 Terra.
  • Gemini 3.7 Flash arrived three weeks after 3.6; Pro cadence is slowing, signaling a two-speed product strategy by tier.

Google's DeepMind team dropped a fresh model card for Gemini 3.7 Flash dated August 13, 2026, and the numbers most buyers will actually price against sit on the coding side. FrontierCode 1.1 moves from 34.4% on Gemini 3.6 Flash to 43.6% on the new one, and the long-context GDM-MRCR v2 8-needle score climbs from 91.8% to 97.0%. DeepSWE v1.1 lands at 65.3%, Terminal-bench 2.1 at 85.8%, and Code Arena Web development at 1588 Elo.

Introductory pricing, running through December 31, 2026, is $0.75 per million input tokens and $3.75 per million output. Standard rates of $1.50 and $7.50 kick in on January 1, 2027, exactly double the intro tier. The context window stays at 1M tokens in and 64K out. Distribution covers Google AI Studio, the Gemini API, Google Antigravity, Gemini Enterprise, the Enterprise Agent Platform, and the Spark tier of the consumer Gemini App. This lands in the middle of a very heavy run of Google coverage this quarter.

DeepMind describes the release as adding "algorithmic improvements to its core reasoning foundation" and "customizable thinking configurations to control the mix of quality, cost and latency," and positions it for "agentic workflows, coding tasks, and enterprise workflows." On safety, the card says the model "performs similarly to Gemini 3.6 Flash across both safety and tone, with low unjustified refusals."

A few gaps to flag. Terminal-bench 3.0, the harder cousin of the 2.1 suite it aces, comes in at just 14.9%, and the card does not explain what accounts for that spread. The knowledge cutoff is March 2026, with some domains capped at January 2025, so anyone building against very recent APIs or events should plan around it. And every number here is Google's own; no independent evaluator has replicated the coding jumps yet.

For teams already shipping coding agents on Flash-class tokens, the pragmatic move is to run your own eval this quarter, while the cheaper rate still applies, and decide whether the January doubling changes unit economics before you commit. The same accounting question is playing out at Anthropic, where session habits are driving Claude Code bills.

What others are reporting

Coverage cluster as of 8h after publish

  1. InfoWorld Read →

    Frames the launch within diverging model-tier economics: Flash commoditizing rapidly while Pro slows, and enterprise governance still blocking agent adoption despite falling prices.

    Gemini 3.7 Flash delivers a noticeably improved developer experience over 3.6 Flash.
  2. The Decoder Read →

    Leads with competitive positioning against Claude Sonnet 5 and GPT-5.6 Terra, and flags the three-week release cadence as the real pricing signal.

    The company calls it its most capable workhorse model yet for coding and AI agents.
  3. DataCamp Read →

    Technical reference: full pricing cliff structure, 1M context window, tunable thinking levels, and benchmark breakdowns across FrontierCode, Code Arena, and DeepSWE.

    Gemini 3.7 Flash is Google's mid-tier workhorse model, positioned as its most capable Flash model for complex coding, agentic workflows, and reliable multi-step execution.