← Back to Insights

Google Ships Three Gemini Flash Models: 17% Fewer Tokens, but No 3.5 Pro

Nils Liu
Google Gemini AI Agents Cybersecurity News

TL;DR

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and the restricted 3.5 Flash Cyber on July 21, 2026. Prices, speed, and benchmarks improved, while 3.5 Pro remains unavailable.

Google Ships Three Gemini Flash Models: 17% Fewer Tokens, but No 3.5 Pro

Whether this release lowers the real cost of AI agents depends on total tokens and tool calls per successful task, not the API price card. If production tests over the next three months show no decline in cost per completed task, the efficiency claim will have failed a practical test. Google supplied the first numbers for that comparison on 2026-07-21, when it released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.

Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. Citing the Artificial Analysis Index, Google says the model uses 17% fewer output tokens than 3.5 Flash. It reports a reduction of as much as 65% under the different conditions used by DataCurve’s DeepSWE benchmark. Those figures cannot be transferred automatically to every workload. Teams still need to include retries, reasoning steps, and external tool calls when calculating the cost of a completed job.

The selected quality benchmarks also improved. Google reports 49% on DeepSWE, compared with 37% for 3.5 Flash; 63.9% versus 49.7% on MLE Bench; and 83.0% versus 78.4% on OSWorld-Verified. These results support the company’s case for better code editing, machine-learning research, and computer use under the stated tests. They remain vendor-selected results, however. TechCrunch independently confirmed the product positioning and release details but did not report a separate reproduction of the scores.

Flash-Lite attaches a price to throughput

Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens. Google cites a speed of 350 output tokens per second from Artificial Analysis and says its Terminal-Bench 2.1 score rose to 54%, from 31% for 3.1 Flash-Lite. The intended workloads include agentic search, document processing, and high-volume subtasks. A company could route low-risk steps to Flash-Lite and reserve a stronger model for review, but the savings materialize only if the additional routing and verification do not create more retries.

Gemini 3.5 Flash Cyber follows a more restricted deployment path. Google paired the specialized model with its CodeMender security agent, where multiple agents detect, validate, and patch vulnerabilities before producing a combined report. Because the capability can be used defensively or misused, Google plans to offer it only to governments and trusted partners in a limited-access pilot. The announcement says the system is competitive at the frontier on CyberGym, but it does not provide a complete independently reproducible result. That limits any conclusion about how much remediation capability the restricted access actually buys.

The missing Pro model matters

The release did not include Gemini 3.5 Pro. TechCrunch notes that Google said in May that Pro was already in internal use and that it expected to roll the model out the following month. By July, product lead Logan Kilpatrick said only that testing with partners continued and that the team hoped the model would land soon. Flash models prioritize speed and cost, while Pro models serve more complex reasoning and coding work; the new releases therefore do not fully fill that gap.

Two outcomes are measurable over the next three to six months: whether 3.6 Flash reduces cost per successful production-agent task by at least 17%, and whether 3.5 Pro becomes broadly available. The first will show whether token savings survive retries and tool expenses. The second will show whether Google can turn flagship-model partner testing into a product customers can actually deploy.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.