deepmind.google web signal

Google Ships Gemini 3.7 Flash, Coding Bench Jumps to 43.6%

5 sources tracking this story
Google Coding Tools ai-business

TL;DR

  • DeepSWE v1.1 score jumped from 48.6% to 65.3% in one generation, a 16.7-point gain Google says surpasses Claude Sonnet 5 on long-horizon repository editing.
  • Introductory pricing of $0.75/$3.75 per million tokens expires December 31, 2026; rates double to $1.50/$7.50 starting January 1, 2027, matching what Gemini 3.6 Flash cost at its own launch.
  • Google shipped three Flash-tier releases in seven weeks; Gemini 3.5 Pro, announced for June 2026, remains in limited partner testing with no public launch date.

Google's DeepMind team dropped a fresh model card for Gemini 3.7 Flash dated August 13, 2026, and the numbers most buyers will actually price against sit on the coding side. FrontierCode 1.1 moves from 34.4% on Gemini 3.6 Flash to 43.6% on the new one, and the long-context GDM-MRCR v2 8-needle score climbs from 91.8% to 97.0%. DeepSWE v1.1 lands at 65.3%, Terminal-bench 2.1 at 85.8%, and Code Arena Web development at 1588 Elo.

Introductory pricing, running through December 31, 2026, is $0.75 per million input tokens and $3.75 per million output. Standard rates of $1.50 and $7.50 kick in on January 1, 2027, exactly double the intro tier. The context window stays at 1M tokens in and 64K out. Distribution covers Google AI Studio, the Gemini API, Google Antigravity, Gemini Enterprise, the Enterprise Agent Platform, and the Spark tier of the consumer Gemini App. This lands in the middle of a very heavy run of Google coverage this quarter.

DeepMind describes the release as adding "algorithmic improvements to its core reasoning foundation" and "customizable thinking configurations to control the mix of quality, cost and latency," and positions it for "agentic workflows, coding tasks, and enterprise workflows." On safety, the card says the model "performs similarly to Gemini 3.6 Flash across both safety and tone, with low unjustified refusals."

A few gaps to flag. Terminal-bench 3.0, the harder cousin of the 2.1 suite it aces, comes in at just 14.9%, and the card does not explain what accounts for that spread. The knowledge cutoff is March 2026, with some domains capped at January 2025, so anyone building against very recent APIs or events should plan around it. And every number here is Google's own; no independent evaluator has replicated the coding jumps yet.

For teams already shipping coding agents on Flash-class tokens, the pragmatic move is to run your own eval this quarter, while the cheaper rate still applies, and decide whether the January doubling changes unit economics before you commit. The same accounting question is playing out at Anthropic, where session habits are driving Claude Code bills.

What others are reporting

Coverage cluster as of 24h after publish

  1. Flags the three-week gap between 3.6 and 3.7 Flash and the continued absence of Gemini 3.5 Pro, framing Flash velocity as evidence of product-line tension inside Google.

    Most intelligent workhorse model yet for coding and agents
  2. DataCamp Read →

    Positions the launch as a direct test of whether Google DeepMind can keep pace with Anthropic and OpenAI while the flagship Pro tier stalls; includes multi-benchmark competitor table.

    The most telling benchmark headline is DeepSWE v1.1, where 3.7 Flash scores 65.3% against 34.4% for its predecessor
  3. FelloAI Read →

    Only source to surface the geographic lockout for EU, UK, Swiss, and Nigerian Spark users, and to benchmark cost directly against GPT-5.6 Luna and DeepSeek V4-Flash.

    Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
  4. Syntax & Signal Read →

    Most granular benchmark table in the set; notes 3.7 Flash is 60 to 70% cheaper than competing frontier models while staying within 1 to 2 points on most evals outside DeepSWE.

    Most intelligent workhorse model yet for coding, agentic reasoning, and complex developer workflows

Shared on Bluesky by 1 AI expert