
Google Releases Three Models Including Gemini 3.6 Flash: 17% Lower Token Costs, Twice as Fast
TechFlow Selected TechFlow Selected

Google Releases Three Models Including Gemini 3.6 Flash: 17% Lower Token Costs, Twice as Fast
These models are designed for the efficiency, latency, and reliability required to build AI Agents at scale.
Author: Google DeepMind
Compiled by: TechFlow
TechFlow Editor's Note: The core focus of the three models released by Google this time is solving real pain points in AI Agent commercialization—cost and speed. 3.6 Flash saves 17% tokens compared to the previous generation, with lower prices but better quality; 3.5 Flash-Lite reaches speeds of 350 tokens/second, making it the fastest 3.5 series model currently; and the specialized cybersecurity model Flash Cyber targets the essential market of code vulnerability repair. This is not a performance benchmark game, but paving the way for large-scale deployment of AI Agents.
Google DeepMind today released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. These models are designed for the efficiency, latency, and reliability required to build AI Agents at scale.
Gemini 3.6 Flash: More Efficient, Higher Quality
3.6 Flash is improved directly based on developer feedback on 3.5 Flash. It not only goes further in coding and knowledge work but also significantly improves token efficiency. According to the Artificial Analysis Index, 3.6 Flash reduces output token consumption by 17% compared to 3.5 Flash. In certain benchmarks such as Datacurve's DeepSWE, the reduction is as high as 65%. Fewer reasoning steps and tool calls are also required to complete multi-step workflows.
This efficiency improvement comes with lower prices. Priced at $1.5 per million input tokens and $7.5 per million output tokens, 3.6 Flash reduces the total cost per Agent task, making building and running Agents more economical.
Even more efficient, 3.6 Flash also outperforms 3.5 Flash in various use cases:
Code editing is more precise, reducing unnecessary modifications and execution loops, reaching 49% on DeepSWE (3.5 Flash is 37%), and significantly improving to 63.9% on the machine learning research benchmark MLE Bench (3.5 Flash is 49.7%)
Computer use capabilities improved, with an OSWorld-Verified score of 83.0% (3.5 Flash is 78.4%). Computer use is now available as a built-in client tool via Gemini API and Gemini Enterprise
Knowledge work performance is better, such as a GDPval-AA v2 benchmark score of 1421 (3.5 Flash is 1349). Customers like Hebbia and Harvey found it particularly excellent in multimodal tasks such as document parsing, chart and data analysis, and report drafting
Customer feedback indicates 3.6 Flash is an improvement in both cost and quality, balancing token efficiency, accuracy, and speed in complex workflows and knowledge-based tasks.
Safety Design
3.6 Flash is equipped with enhanced frontier safety safeguards, covering chemical, biological, radiological, and nuclear (CBRN) as well as cyber attack abuse areas. These safeguards significantly improve the model's resistance to jailbreak attacks. At the same time, the model is trained to minimize refusals for beneficial uses.
Gemini 3.5 Flash-Lite: Built for Scaling Agent Workflows
3.5 Flash-Lite is designed for low-latency tasks and development scenarios requiring high throughput, such as Agent search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. According to Artificial Analysis measurements, it runs at speeds up to 350 output tokens per second. Priced at $0.3 per million input tokens and $2.5 per million output tokens, and with quality significantly superior to 3.1 Flash-Lite, it provides excellent cost-performance for developers and customers running high-volume production tasks.
3.5 Flash-Lite supports efficient scaling of Agent systems. At various thinking levels, the model significantly outperforms 3.1 Flash-Lite. Developers can configure the model based on workload: use minimum and low thinking levels for high-volume tasks, prioritizing low-latency, low-cost execution; or enable higher thinking levels to handle multi-step sub-Agent workloads. The model now also includes computer use as a built-in tool, reliably supporting cross-interface Agent tasks.
It has significant improvements in coding and Agent tasks, such as a Terminal-Bench 2.1 score of 54% (3.1 Flash-Lite is 31%), long context such as GDM-MRCR v2 score of 72.2% (3.1 Flash-Lite is 60.1%), and actual task execution such as GDPval-AA v2 score of 1140 (3.1 Flash-Lite is 642).
In fact, in many Agent and coding evaluations, 3.5 Flash-Lite even surpasses 3 Flash, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%), becoming a faster and stronger choice for 2.5 and 3 Flash workloads.
Early customers highlight the unique combination of speed, intelligence, and cost efficiency of 3.5 Flash-Lite in scaling Agent workflows and data processing tasks.
Gemini 3.5 Flash Cyber in CodeMender: Efficiently Discover and Fix Vulnerabilities
The speed at which AI models discover security vulnerabilities has already surpassed the speed at which current systems fix them. Addressing this growing threat requires a software security approach that is both efficient and powerful.
Flash's performance and efficiency make it an ideal foundation for detecting, validating, and patching code security issues at scale. Gemini 3.5 Flash Cyber is built on 3.5 Flash, fine-tuned specifically for discovering and fixing cybersecurity vulnerabilities, and priced lower per token than large models.
In CodeMender, multiple 3.5 Flash Cyber Agents work together to generate a single comprehensive report, achieving frontier competitive levels on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we have adopted a prudent deployment approach. The model will soon be exclusively available to governments and trusted partners through CodeMender in the form of a limited access pilot project. This will allow frontline defenders to discover and fix critical vulnerabilities before they are exploited, while mitigating broader abuse risks.
Start Using Now
3.6 Flash and 3.5 Flash-Lite are available from today:
Developers can use via Gemini API in Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity
Enterprises can use in the Gemini Enterprise Agent platform. 3.6 Flash is also available in the Gemini Enterprise app
Everyone can use via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search
Feedback is welcome to improve future Gemini models, and we look forward to releasing 3.5 Pro soon.
Join TechFlow official community to stay tuned
Telegram:https://t.me/TechFlowDaily
X (Twitter):https://x.com/TechFlowPost
X (Twitter) EN:https://x.com/BlockFlow_News














