Google releases Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, plus a limited-access Flash Cyber for security teams
Google shipped two new members of its Gemini Flash line on July 21, 2026 — Gemini 3.6 Flash and Gemini 3.5 Flash-Lite — alongside a specialized cybersecurity variant, Gemini 3.5 Flash Cyber, extending the low-cost, high-throughput tier of its model lineup aimed at production agents and high-volume applications.
What's new
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens. Google says it produces 17% fewer output tokens than 3.5 Flash on comparable tasks and shows up to a 65% improvement on the DeepSWE coding-agent benchmark. It's available today across the Gemini API, Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise, and the Gemini app.
Gemini 3.5 Flash-Lite undercuts it further on price: $0.30 per million input tokens and $2.50 per million output tokens. Google cites a throughput of 350 output tokens per second and says it substantially outperforms the prior 3.1 Flash-Lite on agentic workflows. It ships in the Gemini API, Gemini Enterprise, the Gemini app, and Google Search.
Gemini 3.5 Flash Cyber is a narrower release: a Flash model fine-tuned specifically for cybersecurity vulnerability detection, distributed only through Google's CodeMender program as a limited-access pilot for governments and select partners — not a general API release.
As Google put it in its announcement: "Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality."
Context
The Flash tier sits below Google's flagship Pro models and is built for the high-volume, cost-sensitive end of the market — chat assistants, search features, and agentic pipelines that call a model thousands of times per session rather than once. Pairing a faster, cheaper general model (3.6 Flash) with an even cheaper "Lite" variant lets Google offer separate price points depending on whether a workload needs more reasoning headroom or just raw throughput. The Cyber variant is notable for a different reason: it's Google routing a fine-tuned model through an existing program (CodeMender, its automated vulnerability-patching effort) rather than the open API, signaling the model is meant for institutional security use rather than general developer access.
Why it matters
The DeepSWE improvement and lower per-token output count both point at the same competitive pressure across the industry: labs are optimizing for cost-per-completed-task, not just raw capability, since agentic workloads amplify token spend far more than single-turn chat. A 17% cut in output tokens compounds directly into inference cost at scale for anyone running Flash models inside an agent loop. The Flash-Lite price point — under a dollar per million tokens combined — keeps Google aggressive against Flash-tier competition from OpenAI's mini models and Anthropic's Haiku line. The Cyber variant is a smaller signal worth watching: it's an early example of a major lab shipping a security-specialized model through a gated, partner-only channel rather than the public API, a distribution pattern that may become more common as labs look to monetize sensitive capabilities without broad release risk.
Corroborating sources
- Deepmind
https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/
- Blog
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
“Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality.”
- Changelog
https://ai.google.dev/gemini-api/docs/changelog