sakutto
Generative AI

What Is Gemini 3.6 Flash? How It Differs from 3.5 Flash-Lite and Flash Cyber

GeminiGoogleLLM update
What Is Gemini 3.6 Flash? How It Differs from 3.5 Flash-Lite and Flash Cyber

What Is Gemini 3.6 Flash (What's New)

The three models announced this time

3.6 Flash
The small, fast flagship. Balances efficiency and performance and handles most everyday use
3.5 Flash-Lite
An even lighter, cheaper, faster entry-tier model. About 350 tokens/sec
3.5 Flash Cyber
Security-focused. A limited pilot for governments and trusted partners

Gemini 3.6 Flash is the small, fast flagship of the Gemini series, announced by Google in July 2026. Google released three models at once, so first let us establish where each sits and what 3.6 Flash makes new.

The definition of Gemini 3.6 Flash (small, fast flagship)

Gemini 3.6 Flash is the latest in the "Flash" line—the small models that respond quickly and cheaply—of Google's Gemini series. Its biggest advance is that it has been made more efficient, handling the same work with fewer output tokens than the previous 3.5 Flash. Rather than a large, brilliant top-tier model, its role is to process large volumes of everyday tasks with a balance of speed and cost.

View official source →
"According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash" — from Google's official blog

Output tokens are the unit of processing an AI consumes when generating an answer, and they directly affect price and speed. A 17% reduction means the same answer can be produced more cheaply and quickly. For the overall positioning of the Gemini series and the flow from the previous generation, see also the Gemini 3.5 guide.

The three models announced together (3.6 Flash / 3.5 Flash-Lite / 3.5 Flash Cyber)

The point this time is that three models with different characters arrived at once. There is a three-tier structure: the flagship 3.6 Flash, the lighter 3.5 Flash-Lite, and the security-focused, limited-availability 3.5 Flash Cyber. Each has different distribution channels and audiences.

View official source →
"For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. … For enterprises in Gemini Enterprise Agent Platform." — from Google's official blog

3.5 Flash-Lite is an entry-tier model that prioritizes speed and unit cost above all, with about 350 tokens/sec as a highlight. 3.5 Flash Cyber sits in Google's security-research track "CodeMender" and is not available to everyone; it is offered as a limited pilot for government agencies and trusted partners. Because their uses are clearly separated, there is little confusion in choosing even though the names are similar.

Evolution from 3.5 Flash (Efficiency and Performance)

DeepSWE (coding metric) score comparison

Gemini 3.6 Flash49
Gemini 3.5 Flash37

Figures are the DeepSWE (Datacurve) scores from Google's official blog. Bars are on a 0–100 scale.

3.6 Flash's evolution shows on both "efficiency" and "performance." Here we look at the reduction in output tokens and the improvement in coding performance with concrete numbers.

What a 17% output-token reduction means

The headline on efficiency is the reduction in output tokens. According to the Artificial Analysis Index, 3.6 Flash reduces output token usage by 17% compared with 3.5 Flash. When the processing needed to produce the same answer drops, the price falls and the response gets faster.

For services that process a large volume of requests, a 17% reduction directly affects operating cost. Even if the difference per call is small, over hundreds of thousands of calls a month it becomes a difference that cannot be ignored. An efficient small model shows its value precisely in these "handle the volume" situations.

Coding performance (DeepSWE 49% vs 37%)

On performance, there is a clear improvement on coding-type metrics. On DeepSWE, provided by Datacurve, 3.6 Flash records 49% against 3.5 Flash's 37%, meaning the practical quality of code generation has risen.

View official source →
"and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%" — from Google's official blog

Google also says it observes "up to 65% on some benchmarks, such as DeepSWE," but that is a peak value (the best case observed). As a figure you can compare consistently, the improvement from 3.5 Flash's 37% to 3.6 Flash's 49% is the clearer guide. When reading numbers, separating the "maximum" from the "representative" value avoids misunderstanding.

Pricing and Choosing Among the Three Models

Pricing (per million tokens, official)

3.6 Flash
$1.50 input / $7.50 output (down from the previous $9 output)
3.5 Flash-Lite
$0.30 input / $2.50 output; about 350 tokens/sec

Alongside performance, pricing is a key concern. Here we organize the prices of 3.6 Flash and 3.5 Flash-Lite, and how to choose among the three models.

Pricing (3.6 Flash is a price cut)

3.6 Flash lowers the price while raising performance. Pricing is $1.50 input / $7.50 output per million tokens, down from the previous $9 output. The lighter 3.5 Flash-Lite is cheaper still, at $0.30 input / $2.50 output.

Because the efficiency gain (17% fewer output tokens) and the unit-price cut overlap, actual operating cost may fall by more than the headline price difference suggests. In high-volume processing, this double effect adds up.

Choosing among the three models

The three models do not compete; they divide by role. If you prioritize speed and unit cost, choose 3.5 Flash-Lite; for a balance of performance and efficiency across most uses, choose 3.6 Flash; and if you are in scope for security use, choose 3.5 Flash Cyber. There is little to agonize over because their availability is separated to begin with.

General app development, chat, and everyday work such as summarization and classification are well covered by 3.6 Flash. Batch processing where you just want to handle volume and squeeze unit cost suits 3.5 Flash-Lite. Because 3.5 Flash Cyber is limited-availability, for many users it will not even be an option.

Availability and What's Next

Finally, we summarize where you can actually use it and the direction Google has indicated for the future.

Where you can use it (AI Studio, Antigravity, Enterprise)

Availability differs by model. Developers can use it via the Gemini API (Google AI Studio and Android Studio), and 3.6 Flash is also available in Google Antigravity. For enterprises, there is the Gemini Enterprise Agent Platform. To try it first, Google AI Studio, which has a free tier, is the entry point.

A step toward Gemini 4, and conclusion

Google also hinted at the development of a next-generation model this time. It says it has already begun pre-training for Gemini 4, so the Flash-line refresh is also a stepping stone toward the next large model.

View official source →
"We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." — from Google's official blog

In short, Gemini 3.6 Flash advances "cheap, fast, and reasonably smart" by another notch and is well suited to handling the majority of everyday processing. Only when your use leans toward security or the lightest processing do you need to consider Flash Cyber or Flash-Lite. For comparisons with other models and how to choose, see also the Gemini 3.5 guide and the ChatGPT (GPT-5) how-to guide.

Official announcements and technical documents for models are often published in English and take effort to work through. When you want to turn an English web page into a form that is easy to read, the following tool helps.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

FAQ

Q. What is Gemini 3.6 Flash?
It is Google's small, fast flagship Gemini model, announced in July 2026. It cuts output token usage by 17% compared with the previous 3.5 Flash while improving performance on some benchmarks.
Google Official Blog
According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash Google Official Blog
Q. What is the difference between 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber?
3.6 Flash is the small flagship model, 3.5 Flash-Lite is an even lighter, cheaper, faster entry-tier model, and 3.5 Flash Cyber is a security-focused, limited-availability model. Their use cases and availability differ.
Google Official Blog
For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. … For enterprises in Gemini Enterprise Agent Platform. Google Official Blog
Q. Where can I use Gemini 3.6 Flash?
Developers can use it via the Gemini API (Google AI Studio and Android Studio), and 3.6 Flash is also available in Google Antigravity. For enterprises, there is the Gemini Enterprise Agent Platform.
Google Official Blog
For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Google Official Blog

Related Tools

Related Tool Categories

Articles