← Back to Insights

Nano Banana 2 Lite Launch: Google Ships a 4-Second Image Model

Nils Liu
Google DeepMind Nano Banana Gemini AI 影像生成 Generative Media News

TL;DR

Google shipped Nano Banana 2 Lite and Gemini Omni Flash on June 30, pricing images at $0.034 per 1,000. We ran the unit economics on the 4-second generation claim and the math does not close.

Nano Banana 2 Lite Launch: Google Ships a 4-Second Image Model

Google shipped two generative models on June 30: Nano Banana 2 Lite for images, Gemini Omni Flash for video. The listed price is $0.034 per 1,000 images, with generation claimed at four seconds per image. Scale that to a million images and run it against the four-second claim as a GPU-hour estimate, and the numbers stop lining up. If you have a better batch-size or hardware assumption, run the math yourself and come back with what breaks in mine.

What launched: two models, two different production workflows

Nano Banana 2 Lite carries the internal model name gemini-3.1-flash-lite-image and went straight to general availability, reachable through the Gemini API, Google AI Studio, and now baked into Search, the Gemini app, and Photos. Google’s pitch rests on three points: world knowledge for generating scenes and data visualizations that need common-sense context, character consistency across repeated generations, and fast text and localization rendering. Compared with the original Nano Banana (Gemini 2.5 Flash Image), the Google Cloud blog post claims a meaningful jump in image quality and world knowledge, without publishing a benchmark table to back the comparison.

Gemini Omni Flash is in public preview, aimed at video editing rather than generation from scratch. It accepts text, image, and video inputs, and can swap characters, relight scenes, change camera angles, and generate synchronized audio alongside the video. It’s priced at $0.10 per second of output, matching Google’s own Veo 3.1 Fast exactly. Clips are currently capped at 10 seconds, and the input side takes no audio at all, an odd gap for a model carrying the “Omni” name that Google’s materials don’t address.

Both models ship with C2PA content credentials and invisible SynthID watermarking turned on by default, consistent with Google’s approach to generative content over the past two years. Confirmed early adopters include Adobe Firefly, Invideo, WPP’s marketing platform WPP Open, stock library Artlist, Figma (which wired Nano Banana 2 Lite into Figma Weave), and Manus AI, which is testing Omni Flash for real-time autonomous workflows. Google Cloud VP of Product Management Michael Gerstenhaber’s official line, that great creative work moves at the speed of your ideas, reads less like a technical claim and more like a pitch aimed at procurement teams.

What the numbers actually say

Start with the $0.034 figure. Per 1,000 images, that comes to $34 per million, or $0.000034 per image. Google also says a single image takes four seconds end to end. If that four seconds represents one GPU running one inference from start to finish, generating a million images would burn roughly 1,111 GPU-hours. At a conservative $2 an hour for A100-class compute, that’s $2,222 in compute cost against $34 in revenue. On paper, this loses money.

That gap has two possible explanations, and both are plausible. Either the backend batches dozens or hundreds of requests together per GPU cycle, so the real per-image compute time is a fraction of four seconds, or Google is subsidizing the price to drive API volume and lock developers into its ecosystem. Either way, the published four-second latency figure is not a reliable proxy for the underlying compute cost. What actually determines whether this pricing holds up over time is queueing behavior under peak load, a number no third party has measured yet.

Omni Flash’s pricing is worth a second look too. Matching Veo 3.1 Fast at $0.10 per second exactly means Google isn’t undercutting its own product line, it’s segmenting it: Veo 3.1 Fast handles generation from nothing, Omni Flash handles conversational editing on top of existing footage. Same price, different job, which reads more like an effort to deepen lock-in with existing customers than a price play. As for the “state of the art at video editing” claim, that’s currently Google’s own description. No independent lab has run Runway, Kling, or OpenAI’s Sora editing tools against it head to head, so for now that line belongs in the marketing column.

Metrics worth watching over the next few months

Three concrete signals should surface within three to six months. First, whether Adobe Firefly and Figma Weave publish their observed generation latency in production; a meaningful gap during peak hours would confirm just how much batching sits behind that four-second number. Second, whether Google lifts the 10-second cap on Omni Flash or adds audio input, both of which currently rule it out for anything beyond short-form editing. Third, whether OpenAI’s Sora, Runway, or Kling match the $0.10-per-second price point; if none of them do, it suggests Google simply carried over its existing Veo pricing rather than opening a price war.

If this was useful, subscribe to the newsletter for weekly AI PM insights and GenAI case studies.

Related reading: Claude Sonnet 5 Becomes Anthropic Default Model, Nano Banana: Testing Gemini 2.5 Flash’s Image Generation

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.