Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google updates the Flash family

Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google updates the Flash family

Google introduces three new Gemini models: 3.6 Flash more efficient and cheaper than its predecessor, 3.5 Flash-Lite at 350 tokens/s, and 3.5 Flash Cyber dedicated to code security. All optimised for agentic workflows at scale.

On 21 July 2026, Google announced three new models in the Gemini Flash family, designed for those building production AI agents: Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. The stated goal is to offer the best balance of efficiency, speed and quality for agentic workflows at scale.

Gemini 3.6 Flash: better and cheaper

3.6 Flash is the direct upgrade to 3.5 Flash, built on feedback from developers and enterprise customers. The main improvement is not just response quality, but also token efficiency: according to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. On some benchmarks like DeepSWE, savings reach up to 65%.

Pricing drops to $1.50/1M input tokens and $7.50/1M output tokens, lower than 3.5 Flash.

On performance:

  • Coding: DeepSWE goes from 37% to 49%, MLE-Bench from 49.7% to 63.9%
  • Computer use: OSWorld-Verified rises from 78.4% to 83.0%
  • Knowledge work: GDPval-AA v2 goes from 1,349 to 1,421

Figma, Harvey, Hebbia and JetBrains have already adopted 3.6 Flash in their pipelines, reporting tangible improvements on document parsing, data analysis and code migration.

The model ships with enhanced Frontier Safety safeguards against misuse in CBRN and cyber offense domains, with a significant reduction in jailbreak susceptibility.

Gemini 3.5 Flash-Lite: 350 tokens per second

3.5 Flash-Lite is the fastest model in the 3.5 family, optimised for high throughput. Artificial Analysis measures it at 350 output tokens per second. Pricing is $0.30/1M input tokens and $2.50/1M output tokens.

It is not just about speed: compared to the previous generation (3.1 Flash-Lite), 3.5 Flash-Lite makes a substantial quality leap:

  • Terminal-Bench 2.1: 54% vs 31%
  • SWE-Bench Pro: 54.2% vs 49.6% (better than 3 Flash too)
  • OSWorld-Verified: 74.0% vs 65.1%
  • Long context GDM-MRCR v2: 72.2% vs 60.1%

It supports configurable thinking levels (minimal, low, high), native computer use, and is particularly well suited as a subagent in multi-agent systems where 3.6 Flash handles orchestration. Palo Alto Networks, Ramp and other early adopters highlight its value for high-volume data processing and agentic search.

Gemini 3.5 Flash Cyber: code security at scale

3.5 Flash Cyber is a specialised model, fine-tuned on 3.5 Flash to find and fix security vulnerabilities in code. It is deployed within CodeMender, a Google agent that orchestrates multiple Flash Cyber instances in parallel to produce a unified report.

On CyberGym, the reference benchmark for AI-applied cybersecurity, CodeMender + 3.5 Flash Cyber reaches performance competitive with larger frontier models, at a lower cost per token.

The model is not publicly available: it will be accessible only to governments and certified partners through a limited-access pilot programme, aimed at giving defenders a head start before vulnerabilities can be exploited.

Gemini 3.5 Pro: coming soon

Google confirms that Gemini 3.5 Pro is in testing with selected partners and will be available “as soon as it’s ready”. Meanwhile, the team is already working on the pre-training run for Gemini 4, which Google describes as “the most ambitious ever”.

Gemini on AIDeskPro

Stable Gemini models are available on AIDeskPro via Google Cloud with European endpoint (including Milan), with zero data retention and Enterprise agreements. 3.6 Flash and 3.5 Flash-Lite are already under evaluation for integration. As always, they will be added as soon as we have verified they meet the required security guarantees.


Source: Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber - Google , 21 July 2026.