# Google's new Flash trio: cheaper tokens now, Gemini 4 later

> Gemini 3.6 Flash and 3.5 Flash-Lite land in the Gemini app today with lower prices and vendor-reported benchmark jumps — while Gemini 4 has quietly entered pre-training.

_Source: Google announcement via 9to5Google (vendor-reported benchmarks) · 2026-07-21 · 5 min read · Verified against primary sources_

Canonical: https://iyu.app/e/gemini-3-6-flash-launch

## The 60-second version

Google shipped three cheap Gemini models — 3.6 Flash, 3.5 Flash-Lite, and a restricted Flash Cyber — and confirmed Gemini 4 pre-training has begun.

**Key points**

- 3.6 Flash: 17% fewer output tokens (vendor-reported), $1.50/$7.50 per 1M tokens, knowledge cutoff now March 2026
- 3.5 Flash-Lite: $0.30/$2.50 per 1M tokens; Google claims it beats last generation's full-size 3 Flash
- Flash Cyber finds and patches security vulnerabilities — pilot access for governments and trusted partners only
- All benchmarks are Google's own numbers; 3.5 Pro still in partner testing

**Verdict.** A solid efficiency-and-price release if the self-reported numbers hold — but the headline that matters is one sentence: Gemini 4's pre-training run has started.

## Full explainer

> **⚑ Caveat:** Every benchmark and pricing figure below comes from Google's own announcement (as reported by 9to5Google). None of it has been independently verified yet — read the numbers as vendor claims, not measurements.


### What shipped — Three Flash models and one loud hint

While everyone waits for Gemini 3.5 Pro, Google refreshed the small end of the lineup instead: **Gemini 3.6 Flash** (the everyday model, now leaner and cheaper), **Gemini 3.5 Flash-Lite** (the high-throughput tier), and **Gemini 3.5 Flash Cyber** (a locked-down security specialist). And in the same breath, Google confirmed the next frontier run: pre-training for **Gemini 4** has already begun.


### Gemini 3.6 Flash — The pitch is efficiency, not IQ

Google's headline claim for 3.6 Flash is that it does the same work with less: **17% fewer output tokens** than 3.5 Flash (per the Artificial Analysis Index), and fewer reasoning steps and tool calls in multi-step workflows. For agentic use, that's the number that matters — you pay for every intermediate step, so a model that wanders less is a model that costs less, twice over. The sticker price fell too: **$1.50 per million input tokens and $7.50 per million output**, down from $9 on output.

- **-17%** — output tokens vs 3.5 Flash (vendor-reported)
- **$7.50** — per 1M output tokens, down from $9
- **49%** — DeepSWE coding score, up from 37%
- **Mar 2026** — new knowledge cutoff (was Jan 2025)

On capability, Google reports cleaner coding — “higher precision with fewer unwanted code edits and reduced execution loops” — with DeepSWE rising from 37% to 49% and ML-engineering benchmark MLE Bench from 49.7% to 63.9%. Knowledge work (GDPval-AA) climbs from 1349 to 1421, and computer use on OSWorld-Verified goes from 78.4% to 83%. The quietest but maybe most user-visible change: the knowledge cutoff jumps fourteen months, from January 2025 to March 2026.

*Figure: Google's quality-versus-tokens chart for Gemini 3.6 Flash — the efficiency claim in one picture.*


### Gemini 3.5 Flash-Lite — The budget tier catches up to yesterday's mid-tier

Flash-Lite is built for high-throughput, low-latency jobs — agentic search, bulk document processing — at **$0.30 per million input and $2.50 per million output tokens**. Google's comparisons against the March-era 3.1 Flash-Lite show big claimed jumps:

- **Terminal-Bench 2.1 (agentic coding):** 54% vs 31%
- **GDM-MRCR v2 (long context):** 72.2% vs 60.1%
- **GDPval-AA v2 (real-world tasks):** 1140 vs 642
- **Price (input / output, per 1M):** $0.30 / $2.50

The more telling claim: Google says this Lite model now **outperforms the full-size Gemini 3 Flash** of a generation ago — 54.2% vs 49.6% on SWE-Bench Pro and 74.0% vs 65.1% on OSWorld-Verified. If that holds up under independent testing, it's a neat illustration of how quickly capability slides down the price ladder.


### Gemini 3.5 Flash Cyber — A security specialist, deliberately on a leash

The odd one out is Flash Cyber, tuned to detect, validate, and patch code security vulnerabilities at scale — Google's CodeMender tool runs several of these agents together. Google frames it as giving “frontline defenders a head start” on critical vulnerabilities. Note the access model, though: a **limited pilot for governments and trusted partners only**. A model that hunts security holes cheaply is dual-use almost by definition, and the restricted rollout says Google knows it.


### What's next — The small models pay the bills; Gemini 4 is in the oven

Gemini 3.6 Flash and 3.5 Flash-Lite are live in the Gemini app today, with Flash-Lite also headed to Search; developers get both via Google Antigravity, AI Studio, and Android Studio. Gemini 3.5 Pro remains “testing with partners,” arriving “as soon as it's ready.” And the sentence that will outlive this whole announcement: DeepMind says it has started its most ambitious pre-training run yet — for Gemini 4.

> Token efficiency is the real price of agents — and the durable news here is one sentence about Gemini 4.


## Primary sources

- [9to5Google — Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4](https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/)

---
_Published by iyu (https://iyu.app) — the day's AI news, checked against primary sources and rewritten in plain language. Free to quote with attribution and a link to the canonical URL._
