Google Ships Three Gemini Flash Models: 17% Fewer Tokens, but No 3.5 Pro
TL;DR
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and the restricted 3.5 Flash Cyber on July 21, 2026. Prices, speed, and benchmarks improved, while 3.5 Pro remains unavailable.
Whether this release lowers the real cost of AI agents depends on total tokens and tool calls per successful task, not the API price card. If production tests over the next three months show no decline in cost per completed task, the efficiency claim will have failed a practical test. Google supplied the first numbers for that comparison on 2026-07-21, when it released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.
Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. Citing the Artificial Analysis Index, Google says the model uses 17% fewer output tokens than 3.5 Flash. It reports a reduction of as much as 65% under the different conditions used by DataCurve’s DeepSWE benchmark. Those figures cannot be transferred automatically to every workload. Teams still need to include retries, reasoning steps, and external tool calls when calculating the cost of a completed job.
The selected quality benchmarks also improved. Google reports 49% on DeepSWE, compared with 37% for 3.5 Flash; 63.9% versus 49.7% on MLE Bench; and 83.0% versus 78.4% on OSWorld-Verified. These results support the company’s case for better code editing, machine-learning research, and computer use under the stated tests. They remain vendor-selected results, however. TechCrunch independently confirmed the product positioning and release details but did not report a separate reproduction of the scores.
Flash-Lite attaches a price to throughput
Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens. Google cites a speed of 350 output tokens per second from Artificial Analysis and says its Terminal-Bench 2.1 score rose to 54%, from 31% for 3.1 Flash-Lite. The intended workloads include agentic search, document processing, and high-volume subtasks. A company could route low-risk steps to Flash-Lite and reserve a stronger model for review, but the savings materialize only if the additional routing and verification do not create more retries.
Gemini 3.5 Flash Cyber follows a more restricted deployment path. Google paired the specialized model with its CodeMender security agent, where multiple agents detect, validate, and patch vulnerabilities before producing a combined report. Because the capability can be used defensively or misused, Google plans to offer it only to governments and trusted partners in a limited-access pilot. The announcement says the system is competitive at the frontier on CyberGym, but it does not provide a complete independently reproducible result. That limits any conclusion about how much remediation capability the restricted access actually buys.
The missing Pro model matters
The release did not include Gemini 3.5 Pro. TechCrunch notes that Google said in May that Pro was already in internal use and that it expected to roll the model out the following month. By July, product lead Logan Kilpatrick said only that testing with partners continued and that the team hoped the model would land soon. Flash models prioritize speed and cost, while Pro models serve more complex reasoning and coding work; the new releases therefore do not fully fill that gap.
Two outcomes are measurable over the next three to six months: whether 3.6 Flash reduces cost per successful production-agent task by at least 17%, and whether 3.5 Pro becomes broadly available. The first will show whether token savings survive retries and tool expenses. The second will show whether Google can turn flagship-model partner testing into a product customers can actually deploy.
Sources:
Related Articles
Google Launches Gemini 3.7 Flash: Half-Price Until Year-End, With Agent Costs Still Tied to Retry Rates
Google positions Gemini 3.7 Flash as a workhorse for coding and AI agents, with higher vendor benchmarks and temporary half-price access, while architecture, training methods, and production retry rates remain undisclosed.
AI Agents Breach Taiwan Government Systems in Four Days, Leaving 12 Attack Waves
Dream reconstructed 12 waves of a multi-agent intrusion from a 160 MB workspace, while Taiwan confirmed an overseas AI-assisted attack on government agencies in July.