Skip to content
← Back to Insights

Google Launches Gemini 3.8 Flash: Token Costs and Cyber Access Behind the Low Price

GenAI News Gemini Cybersecurity

TL;DR

Google launched Gemini 3.8 Flash and a Flash Cyber variant restricted to trusted defenders; introductory unit prices stay low, but longer reasoning can raise the actual bill.

Google Launches Gemini 3.8 Flash: Token Costs and Cyber Access Behind the Low Price

Whether Gemini 3.8 Flash reduces the total cost of an agent workflow cannot be determined from its API price alone. A practical three-month test would compare the change in task success with the change in average output tokens and tool calls. If token use or tool calls rise faster than successful completions, the low-cost thesis fails. Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on 2026-09-02. The general model is already available through developer, enterprise, and consumer products, while the cyber version is restricted to vetted defenders.

Google positions the general model for software engineering, agentic work, and multi-step reasoning in specialized fields. Its introductory price is $0.75 per million input tokens and $3.75 per million output tokens, unchanged from Gemini 3.7 Flash, which arrived three weeks earlier. However, Google explicitly says that 3.8 performs additional reasoning steps and calls tools iteratively on complex jobs. At higher effort settings, it may consume more tokens to maximize performance. Developers who prioritize efficiency can lower the effort level or continue using the still-supported 3.7 Flash. As The Verge notes, an unchanged unit price therefore does not guarantee an unchanged bill per completed task.

The release makes usage volume as important as benchmark performance. Google reports 54.9 percent on HLE-Verified and says the model improves on 3.7 Flash in DeepSWE v1.1 as well as finance and legal agent evaluations. These figures are mostly vendor-reported. The announcement does not disclose average token use, latency, or total task cost at each effort setting, so a buyer cannot turn the leaderboard results directly into a production budget. The introductory price also expires on 2026-12-31. From 2027-01-01, Google says prices will rise to $1.50 per million input tokens and $7.50 per million output tokens.

The access boundary around Flash Cyber

Gemini 3.8 Flash Cyber uses the same underlying intelligence but is optimized for vulnerability discovery and automated patching. Google says the model exceeded a 70 percent success rate on an internal test spanning complex codebases in 20 programming languages. On the external CWE-Bench patching benchmark, it recorded a pass@1 of 47.2 percent, close to 47.8 percent for another leading frontier model. Google also says its Chrome Security team obtained 2.6 times as many correct vulnerability patches as it did from much larger commercial models. Those real-world examples come from Google and named partners; an outside group has not yet published a full reproduction of the test conditions.

Distribution is deliberately narrow. The standard model includes safeguards covering chemical, biological, radiological, nuclear, and offensive cyber misuse under Google’s Frontier Safety Framework. The Cyber variant has more permissive cybersecurity mitigations and is available only through the Fairwind Program. Google names government authorities, critical-infrastructure operators, and software maintainers as intended recipients. Most developers can test the standard model today, but they cannot independently verify Flash Cyber’s patching claims or compare its false-positive rate without approval.

For the next three to six months, procurement teams can run a fixed set of jobs on versions 3.7 and 3.8 and record success rate, total tokens, tool calls, and latency. They should also recalculate the results with the 2027 list price rather than only the launch promotion. For Flash Cyber, the observable variables are how many external organizations receive Fairwind access and whether independent vulnerability and patching reproductions appear. Google’s announcement does not yet provide either figure, leaving the breadth and repeatability of deployment unresolved.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.