Gemini 3.6 Flash is now available as Google expands its artificial intelligence lineup with two general-purpose models and a specialized cybersecurity system.
The company launched Gemini 3.6 Flash alongside Gemini 3.5 Flash-Lite on July 21, positioning the models for developers and businesses building high-volume AI agents. It also announced Gemini 3.5 Flash Cyber, a security-focused model that will initially be restricted to selected governments and trusted partners.
Google says the releases are designed to offer different combinations of intelligence, speed and operating cost. Gemini 3.6 Flash targets complex coding, knowledge work and multimodal analysis, while Flash-Lite prioritizes low latency and high throughput.
Gemini 3.6 Flash targets production AI agents

Google describes Gemini 3.6 Flash as its new workhorse model for agentic and multimodal workloads.
The model is designed to complete coding projects, analyze documents and charts, conduct multi-step research and interact with software tools. Google says it requires fewer reasoning steps, conversational turns and tool calls than Gemini 3.5 Flash when completing complex workflows.
That efficiency could matter for businesses operating AI agents at scale because every generated token and external tool call can increase costs and processing time.
Google reports that Gemini 3.6 Flash consumed 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index. On some specialized software-engineering evaluations, the company observed substantially larger reductions, although benchmark performance may not always reflect results in production applications.
The model supports text, images, video, audio and PDF files as inputs, while producing text output. It offers an input context limit of more than one million tokens and a maximum output of 65,536 tokens.
Coding and multimodal performance improve
Coding is one of the main areas Google is emphasizing with Gemini 3.6 Flash.
The company says the model generates production-ready code with fewer unnecessary file changes and fewer repeated debugging loops. It is also designed to inspect projects programmatically before attempting changes, which may improve accuracy on complex migration or diagnostic tasks.
Google reported a score of 49% for Gemini 3.6 Flash on the DeepSWE software-engineering benchmark, compared with 37% for Gemini 3.5 Flash. It also cited improvements in machine-learning research, computer interaction and knowledge-work evaluations.
The model’s multimodal capabilities target tasks such as parsing reports, extracting information from documents, interpreting charts and converting visual plans into structured outputs.
Google acknowledged one trade-off in its developer guidance: human evaluators sometimes preferred earlier models for visual styling and page layout. Developers may need to provide more detailed design instructions when using Gemini 3.6 Flash to create user interfaces.
Google cuts Gemini 3.6 Flash output pricing
Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens through the Gemini API’s standard paid tier.
The input price is unchanged from Gemini 3.5 Flash, while the output price is lower than the earlier model’s $9 per million tokens. Google argues that the reduced price, combined with lower token consumption, should decrease the total cost of running agentic workflows.
The model is generally available for production use through the Gemini API, Google AI Studio and Android Studio. It is also available through Google’s enterprise AI products and the consumer Gemini application.
Gemini 3.6 Flash has become the default model for Google’s Antigravity managed agent, although developers can select another supported model.
Gemini 3.5 Flash-Lite prioritizes speed
Gemini 3.5 Flash-Lite is intended for applications where speed, volume and cost matter more than maximum reasoning capability.

Google says it is the fastest model in the Gemini 3.5 family, producing about 350 output tokens per second in testing by Artificial Analysis. Its target uses include document extraction, agent-based search, classification, structured data processing and workloads involving large numbers of relatively short tasks.
The model costs $0.30 per million input tokens and $2.50 per million output tokens.
Like Gemini 3.6 Flash, Flash-Lite supports a one-million-token context window, a maximum output of 64,000 tokens, configurable thinking and built-in tools including computer use.
Developers can select different thinking levels depending on the workload. The minimal setting prioritizes speed and low cost, while medium or high settings allocate more processing to planning, tool use and multi-step reasoning.
Google says Flash-Lite significantly outperforms Gemini 3.1 Flash-Lite on coding, long-context and real-world task evaluations. It is generally available through Google’s developer and enterprise platforms and is also being introduced in Google Search.
Flash Cyber receives restricted release
Gemini 3.5 Flash Cyber is a specialized version of Gemini 3.5 Flash trained to identify, validate and help repair software vulnerabilities.
Google is pairing the model with CodeMender, its AI-driven code-security agent. The system can coordinate multiple Flash Cyber agents and combine their findings into one report or proposed remediation process.
The company says the model is cheaper per token than larger frontier systems while delivering competitive results on the CyberGym security benchmark. Google has not published general API pricing for Flash Cyber because it will not initially be offered as a broadly available model.
Access will instead be limited to a small number of governments and trusted partners through a pilot program. Google said the cautious deployment reflects the dual-use nature of cybersecurity technology: a system capable of finding vulnerabilities could help defenders, but similar capabilities might also be misused.
Google plans to expand access over time but has not provided a public timetable.

Google divides AI workloads by cost and risk
The three releases illustrate Google’s increasingly segmented AI strategy.
Gemini 3.6 Flash addresses demanding general-purpose workloads without moving into the highest-cost flagship category. Flash-Lite targets developers processing documents, searches and structured data at scale. Flash Cyber applies similar efficiency principles to a sensitive field where Google believes unrestricted access would create additional risks.
Google also confirmed that Gemini 3.5 Pro remains in testing with selected partners. Separately, the company said it has begun pretraining Gemini 4, although it did not provide a release date or detailed specifications.
The next test will be whether developers see the same efficiency gains in real applications that Google reports in benchmarks. Adoption will depend not only on model quality, but also on reliability, migration requirements, tool compatibility and the total cost of completing each task.
Read Also: Xiaomi Mix Fold 5, Redmi Note 17 Gain Certifications






