Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Jul 21, 2026
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models:
- 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.
- 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.
Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
3.6 Flash: More efficient and better quality than 3.5 Flash
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency is also combined with a lower price than 3.5 Flash. At 7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases:
- Higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and significant improvement in ML Research on MLE Bench (63.9% vs. 49.7%).
- Improved computer use capabilities on OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
- Outperforms 3.5 Flash in knowledge work (GDPval-AA v2: 1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks.
3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses.
3.5 Flash-Lite: Built to scale agentic workflows
Gemini 3.5 Flash-Lite is designed for both low-latency tasks and tasks where high throughput is critical, like agentic search and document processing. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at 2.5/1M output tokens, with significantly better quality than 3.1 Flash-Lite.
On many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%).
3.5 Flash Cyber in CodeMender
Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models. Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together, it reaches competitive performance at the frontier on CyberGym.
Given the dual-use nature of this technology, the model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
Availability
3.6 Flash and 3.5 Flash-Lite are available starting today:
- For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity.
- For enterprises in Gemini Enterprise Agent Platform. 3.6 Flash is also available in the Gemini Enterprise app.
- For everyone via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.