Author: Google DeepMind
Compiled by: Deep Tide TechFlow
Deep Tide Guide: The core of the three models released by Google this time is to solve the real pain points of commercializing AI Agents—cost and speed. 3.6 Flash saves 17% in token costs compared to the previous generation, with a lower price but better quality; 3.5 Flash-Lite reaches a speed of 350 tokens/second, making it the fastest in the 3.5 series; while the dedicated cybersecurity model Flash Cyber targets the essential market of code vulnerability fixing. This is not a performance benchmarking game, but a path to large-scale deployment of AI Agents.
Google DeepMind today released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. These models are designed for the efficiency, latency, and reliability required for large-scale construction of AI Agents.
Gemini 3.6 Flash: More Efficient, Higher Quality
3.6 Flash is directly improved based on developer feedback on 3.5 Flash. It not only makes further advancements in coding and knowledge work but also significantly enhances token efficiency. According to the Artificial Analysis Index, 3.6 Flash reduces output token consumption by 17% compared to 3.5 Flash. In certain benchmarks like Datacurve's DeepSWE, the reduction reaches up to 65%. The reasoning steps and tool calls required to complete multi-step workflows are also fewer.
This efficiency improvement comes with lower prices. Priced at $1.5 per million input tokens and $7.5 per million output tokens, 3.6 Flash reduces the total cost per Agent task, making the construction and operation of Agents more economical.
Despite being more efficient, 3.6 Flash outperforms 3.5 Flash in various use cases:
Code editing is more precise, reducing unnecessary modifications and execution loops, reaching 49% on DeepSWE (37% for 3.5 Flash), and significantly improving to 63.9% on the machine learning research benchmark MLE Bench (49.7% for 3.5 Flash).
Computer operation capability has improved, with an OSWorld-Verified score of 83.0% (78.4% for 3.5 Flash). Computer operation is now provided as a built-in client tool through the Gemini API and Gemini Enterprise.
Knowledge work performance is better, scoring 1421 on the GDPval-AA v2 benchmark (1349 for 3.5 Flash). Clients such as Hebbia and Harvey have found it particularly outstanding in multimodal tasks like document parsing, chart and data analysis, and report drafting.
Customer feedback indicates that 3.6 Flash represents improvements in cost and quality, balancing token efficiency, accuracy, and speed in complex workflows and knowledge-based tasks.
Security Design
3.6 Flash comes equipped with enhanced cutting-edge security measures, covering chemical, biological, radiological and nuclear (CBRN) threats and cyberattack misuse areas. These protections significantly increase the model's resistance to jailbreaking attacks. At the same time, the model is trained to minimize refusals for beneficial uses.
Gemini 3.5 Flash-Lite: Born for Scaling Agent Workflows
3.5 Flash-Lite is designed for low-latency tasks and development scenarios requiring high throughput, such as Agent searches and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. According to Artificial Analysis measurements, it runs at a speed of 350 output tokens per second. Priced at $0.3 per million input tokens and $2.5 per million output tokens, it offers significantly better quality than 3.1 Flash-Lite, providing excellent cost-effectiveness for developers and customers running high-volume production tasks.
3.5 Flash-Lite supports efficient scaling of Agent systems. At all levels of thinking, this model significantly outperforms 3.1 Flash-Lite. Developers can configure the model according to workload: using minimal and low thinking levels for large-batch tasks, prioritizing low-latency, low-cost execution; or enabling higher thinking levels to handle multi-step sub-Agent workloads. The model now also includes computer operations as a built-in tool, reliably supporting cross-interface Agent tasks.
It has significant improvements in coding and Agent tasks, such as scoring 54% on Terminal-Bench 2.1 (31% for 3.1 Flash-Lite), achieving 72.2% on long contexts like GDM-MRCR v2 (60.1% for 3.1 Flash-Lite), and scoring 1140 in actual task execution on GDPval-AA v2 (642 for 3.1 Flash-Lite).
In fact, in many Agent and coding evaluations, 3.5 Flash-Lite even surpasses 3 Flash, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%), becoming a faster and stronger choice for 2.5 and 3 Flash workloads.
Early customers emphasized the unique combination of speed, intelligence, and cost efficiency of 3.5 Flash-Lite in scaling Agent workflows and data processing tasks.
Gemini 3.5 Flash Cyber in CodeMender: Efficiently Discovering and Fixing Vulnerabilities
The speed at which AI models discover security vulnerabilities has outpaced the current systems that fix them. Addressing this growing threat requires a software security approach that is both efficient and powerful.
The performance and efficiency of Flash make it an ideal foundation for large-scale detection, validation, and remediation of code security issues. Gemini 3.5 Flash Cyber is built on 3.5 Flash, fine-tuned specifically for discovering and fixing cybersecurity vulnerabilities, and is priced lower than large models per token.
In CodeMender, multiple 3.5 Flash Cyber Agents work together to generate a single comprehensive report, reaching cutting-edge competitive levels on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we are taking a cautious deployment approach. This model will soon be offered via CodeMender in the form of a restricted access pilot project specifically for governments and trusted partners. This will allow frontline defenders to detect and fix critical vulnerabilities before they can be exploited, while also mitigating broader abuse risks.
Get Started Immediately
3.6 Flash and 3.5 Flash-Lite are available starting today:
Developers can use them through the Gemini API in Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity.
Enterprises can use them in the Gemini Enterprise Agent platform. 3.6 Flash is also available in the Gemini Enterprise application.
Everyone can use them through the Gemini app. 3.5 Flash-Lite is also being rolled out in Google Search.
We welcome feedback to improve future Gemini models, and we look forward to releasing 3.5 Pro soon.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。