Google's new Flash trio: cheaper tokens now, Gemini 4 later

Gemini 3.6 Flash and 3.5 Flash-Lite land in the Gemini app today with lower prices and vendor-reported benchmark jumps — while Gemini 4 has quietly entered pre-training.

✓ Verified Source Google announcement via 9to5Google (vendor-reported benchmarks) ⚑ Model release

The 60-second version

Google shipped three cheap Gemini models — 3.6 Flash, 3.5 Flash-Lite, and a restricted Flash Cyber — and confirmed Gemini 4 pre-training has begun.

Key points

  • 3.6 Flash: 17% fewer output tokens (vendor-reported), $1.50/$7.50 per 1M tokens, knowledge cutoff now March 2026
  • 3.5 Flash-Lite: $0.30/$2.50 per 1M tokens; Google claims it beats last generation's full-size 3 Flash
  • Flash Cyber finds and patches security vulnerabilities — pilot access for governments and trusted partners only
  • All benchmarks are Google's own numbers; 3.5 Pro still in partner testing

Verdict. A solid efficiency-and-price release if the self-reported numbers hold — but the headline that matters is one sentence: Gemini 4's pre-training run has started.

What shippedThree Flash models and one loud hint

While everyone waits for Gemini 3.5 Pro, Google refreshed the small end of the lineup instead: Gemini 3.6 Flash (the everyday model, now leaner and cheaper), Gemini 3.5 Flash-Lite (the high-throughput tier), and Gemini 3.5 Flash Cyber (a locked-down security specialist). And in the same breath, Google confirmed the next frontier run: pre-training for Gemini 4 has already begun.

Gemini 3.6 FlashThe pitch is efficiency, not IQ

Google's headline claim for 3.6 Flash is that it does the same work with less: 17% fewer output tokens than 3.5 Flash (per the Artificial Analysis Index), and fewer reasoning steps and tool calls in multi-step workflows. For agentic use, that's the number that matters — you pay for every intermediate step, so a model that wanders less is a model that costs less, twice over. The sticker price fell too: $1.50 per million input tokens and $7.50 per million output, down from $9 on output.

-17%output tokens vs 3.5 Flash (vendor-reported)
$7.50per 1M output tokens, down from $9
49%DeepSWE coding score, up from 37%
Mar 2026new knowledge cutoff (was Jan 2025)

On capability, Google reports cleaner coding — “higher precision with fewer unwanted code edits and reduced execution loops” — with DeepSWE rising from 37% to 49% and ML-engineering benchmark MLE Bench from 49.7% to 63.9%. Knowledge work (GDPval-AA) climbs from 1349 to 1421, and computer use on OSWorld-Verified goes from 78.4% to 83%. The quietest but maybe most user-visible change: the knowledge cutoff jumps fourteen months, from January 2025 to March 2026.

Google's quality-versus-tokens chart for Gemini 3.6 Flash — the efficiency claim in one picture.
Google's quality-versus-tokens chart for Gemini 3.6 Flash — the efficiency claim in one picture. · Google via 9to5Google

Gemini 3.5 Flash-LiteThe budget tier catches up to yesterday's mid-tier

Flash-Lite is built for high-throughput, low-latency jobs — agentic search, bulk document processing — at $0.30 per million input and $2.50 per million output tokens. Google's comparisons against the March-era 3.1 Flash-Lite show big claimed jumps:

Terminal-Bench 2.1 (agentic coding)54% vs 31%
GDM-MRCR v2 (long context)72.2% vs 60.1%
GDPval-AA v2 (real-world tasks)1140 vs 642
Price (input / output, per 1M)$0.30 / $2.50

The more telling claim: Google says this Lite model now outperforms the full-size Gemini 3 Flash of a generation ago — 54.2% vs 49.6% on SWE-Bench Pro and 74.0% vs 65.1% on OSWorld-Verified. If that holds up under independent testing, it's a neat illustration of how quickly capability slides down the price ladder.

Gemini 3.5 Flash CyberA security specialist, deliberately on a leash

The odd one out is Flash Cyber, tuned to detect, validate, and patch code security vulnerabilities at scale — Google's CodeMender tool runs several of these agents together. Google frames it as giving “frontline defenders a head start” on critical vulnerabilities. Note the access model, though: a limited pilot for governments and trusted partners only. A model that hunts security holes cheaply is dual-use almost by definition, and the restricted rollout says Google knows it.

What's nextThe small models pay the bills; Gemini 4 is in the oven

Gemini 3.6 Flash and 3.5 Flash-Lite are live in the Gemini app today, with Flash-Lite also headed to Search; developers get both via Google Antigravity, AI Studio, and Android Studio. Gemini 3.5 Pro remains “testing with partners,” arriving “as soon as it's ready.” And the sentence that will outlive this whole announcement: DeepMind says it has started its most ambitious pre-training run yet — for Gemini 4.

Token efficiency is the real price of agents — and the durable news here is one sentence about Gemini 4.