deepmind.google web signal

Google Ships Gemini 3.7 Flash, Coding Bench Jumps to 43.6%

Google Coding Tools ai-business

TL;DR

  • Gemini 3.7 Flash launched August 13, 2026, keeping a 1M-token input window and 64K output cap.
  • Introductory pricing runs $0.75 input and $3.75 output per million tokens through December 31, 2026, doubling to $1.50 and $7.50 on January 1, 2027.
  • FrontierCode 1.1 climbs from 34.4% on Gemini 3.6 Flash to 43.6%, and GDM-MRCR v2 8-needle rises from 91.8% to 97.0%.

Google's DeepMind team dropped a fresh model card for Gemini 3.7 Flash dated August 13, 2026, and the numbers most buyers will actually price against sit on the coding side. FrontierCode 1.1 moves from 34.4% on Gemini 3.6 Flash to 43.6% on the new one, and the long-context GDM-MRCR v2 8-needle score climbs from 91.8% to 97.0%. DeepSWE v1.1 lands at 65.3%, Terminal-bench 2.1 at 85.8%, and Code Arena Web development at 1588 Elo.

Introductory pricing, running through December 31, 2026, is $0.75 per million input tokens and $3.75 per million output. Standard rates of $1.50 and $7.50 kick in on January 1, 2027, exactly double the intro tier. The context window stays at 1M tokens in and 64K out. Distribution covers Google AI Studio, the Gemini API, Google Antigravity, Gemini Enterprise, the Enterprise Agent Platform, and the Spark tier of the consumer Gemini App. This lands in the middle of a very heavy run of Google coverage this quarter.

DeepMind describes the release as adding "algorithmic improvements to its core reasoning foundation" and "customizable thinking configurations to control the mix of quality, cost and latency," and positions it for "agentic workflows, coding tasks, and enterprise workflows." On safety, the card says the model "performs similarly to Gemini 3.6 Flash across both safety and tone, with low unjustified refusals."

A few gaps to flag. Terminal-bench 3.0, the harder cousin of the 2.1 suite it aces, comes in at just 14.9%, and the card does not explain what accounts for that spread. The knowledge cutoff is March 2026, with some domains capped at January 2025, so anyone building against very recent APIs or events should plan around it. And every number here is Google's own; no independent evaluator has replicated the coding jumps yet.

For teams already shipping coding agents on Flash-class tokens, the pragmatic move is to run your own eval this quarter, while the cheaper rate still applies, and decide whether the January doubling changes unit economics before you commit. The same accounting question is playing out at Anthropic, where session habits are driving Claude Code bills.