9to5google.com web signal

Google launches Gemini 3.6 Flash, teases Gemini 4 pre-training

5 sources tracking this story

TL;DR

  • Gemini 3.6 Flash cuts output tokens 17% versus 3.5 Flash, reaching 65% reduction on DeepSWE benchmarks, compressing enterprise costs without requiring a model upgrade.
  • 3.5 Flash-Lite runs at 350 output tokens per second at $0.30 per million input tokens, targeting latency-sensitive, high-throughput workloads competitors have not directly addressed at this price.
  • Flash Cyber is gated to governments and trusted partners via the CodeMender pilot; Google declined to make offense-capable AI available as a self-serve API.

Google's latest Gemini release, reported by 9to5Google, is less a headline model moment and more a repricing of the middle of the stack. Gemini 3.6 Flash lands at $1.50 per million input tokens and $7.50 per million output, with a cheaper 3.5 Flash-Lite tier at $0.30 and $2.50. On top of the sticker, Google says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on comparable multi-step work, which is the real price cut hiding inside the price cut.

The benchmark numbers the company chose to emphasize are coding and computer-use ones. DeepSWE moves from 37% to 49%, and the knowledge cutoff jumps from January 2025 to March 2026. For anyone wiring an assistant around current libraries, SDKs or documentation, that cutoff shift is a bigger practical unlock than most single-digit benchmark deltas.

The more interesting release is Gemini 3.5 Flash Cyber, a security-tuned variant Google is deliberately not opening up. It ships as a limited-access pilot for governments and trusted partners, and powers CodeMender, Google's automated vulnerability-patching agent. According to SiliconANGLE, early enterprise testers of CodeMender include Salesforce, Robinhood and Palo Alto Networks, while Harvey and Hebbia are cited as early users of 3.6 Flash. The gating is the notable posture: a hyperscaler that usually leads with "generally available" is drawing a line around an offense-capable tool.

Buried near the end is the sentence worth circling. Google says it has "already started our most ambitious pre-training run yet, for Gemini 4," while Gemini 3.5 Pro is still in partner testing. Take that as reported, not delivered. The honest caveat is that the reporting doesn't give a Gemini 4 timeline, doesn't quantify Cyber's lift on real-world CVEs, and doesn't say how CodeMender's fixes are being validated once they land in production code.

The upside case is straightforward. If the Flash tier keeps compressing on price and efficiency, it pressures OpenAI and Anthropic at the workhorse layer where most enterprise inference spend actually sits, and gives builders on Harvey- and Hebbia-style stacks a cheaper substrate to iterate on while everyone waits to see what Gemini 4 actually is.

What others are reporting

Coverage cluster as of 2h after publish

  1. Google Blog Read →

    First-party source with full pricing ($1.50/$7.50 per 1M tokens for 3.6 Flash; $0.30/$2.50 for Flash-Lite), CodeMender pilot restrictions, and the Gemini 4 pre-training statement.

    3.6 Flash reduces output token usage by 17% compared to 3.5 Flash, all at a lower cost per output token.
  2. The Decoder Read →

    Names rivals (GPT-5.6 Sol, Anthropic Fable/Mythos, Kimi K3, GLM-5.2), reports Gemini 3.5 Pro is months behind due to coding issues, and frames the Gemini 4 announcement as damage control.

    As long as 3.5 Pro stays in private testing, Google doesn't have a public model that competes at the top.
  3. The Next Web Read →

    Frames Flash Cyber as a direct price-based counter to Anthropic's Mythos security model and reads the full release as a defensive market-flooding move while Google's flagship stays absent.

  4. Thurrott Read →

    Provides platform distribution detail (Gemini app, Google Search, AI Studio, Android Studio, Enterprise) and frames the releases as compensating for the Gemini 3.5 Pro delay first announced at I/O.

    3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency.