Google Launches Gemini 3.8 Live: Spoken Progress Needs to Help the User
TL;DR
Gemini 3.8 Live and its extended-thinking version reach general availability in the API with 97 languages. Background reasoning can reduce silence, but useful updates, task status and audio costs determine product value.
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 2026-09-15, allowing a voice model to keep speaking while reasoning in the background. Public evidence does not yet establish whether updates during a wait reduce repeated user commands. That gap affects a product decision: an assistant can leave less silence, but users still need to know whether their original request is being processed.
The announcement and API release notes carry a September 15 date without an exact activation time. SiliconANGLE displays an update at 19:41 EDT that day, equivalent to 07:41 on September 16 in Taipei. This article is dated September 16 in Taipei and covers newly available capabilities announced under the previous day’s date. The news update is not treated as the model’s activation time.
Google lists both models as generally available in the Gemini API and offers access through AI Studio. The Live API interface documentation still labels the service Preview, distinct from model availability. Gemini Enterprise remains a private preview, with its customer-experience offering coming later. Consumers can access the standard model through Search Live; Extended Thinking starts rolling into Gemini Live and subscription-dependent Workspace features. The models support 97 languages and switching languages within a conversation. The sources provide no independent results for Taiwanese accents or Traditional Chinese customer service.
Extended Thinking can continue reasoning after a spoken response and report the result later. The API documentation warns that clients must keep listening for messages after receiving turnComplete; interaction_status distinguishes ongoing processing from an idle session. A developer who retains a workflow that ends when speech stops could close an interaction while the model is still retrieving information.
Google demonstrates multistep bookings and turning sketches plus spoken feedback into React components, without disclosing the scale of production customer use. Its announcement compiles scores of 68.6% on τ-Voice and 35.1% on Sierra’s banking tasks for Extended Thinking. The methodology specifies high reasoning, default sampling and, unless noted otherwise, a single attempt. Artificial Analysis conducted its own tests. Scores from different task sets are not directly comparable and do not represent success rates among real customers.
Let updates follow changes in the search results
Suppose an assistant is comparing hotels. Confirming that it has received the dates and budget can reassure the user that it heard correctly. Briefly explaining that available options fail a requirement could help the user revise the request early. Repeatedly announcing that it is still searching may instead interrupt the user without adding information for a decision. This is a design judgment derived from background reasoning, not a measured outcome of the demonstrations.
I would tie spoken updates to useful changes in status and retain a client-side option to wait quietly. The model’s proactive audio cannot currently be disabled; controlling playback should not be assumed to stop generation charges. A product need not narrate every internal reasoning step when someone simply wants an answer. It can invite participation when more information or confirmation is needed. Clear distinctions between receipt, ongoing retrieval and completion also help prevent fluent speech from implying that a transaction has already been carried out.
Narration length affects charges as well. The pricing page lists audio input at $3 per million tokens and audio output at $12, approximately $0.005 per input minute and $0.018 per output minute. These are separate audio-usage rates: a minute-long call is not automatically a minute of generated speech. Text and thinking tokens have separate pricing, and other system costs are excluded. Filling a wait with more narration could increase audio spending; whether it reduces repeated requests still needs verification.
For customer-service products, API availability supplies a tool for changing the waiting experience. I would treat progress updates as a way to help someone continue the same task, retaining search history and confirmed requirements so that an interruption can revise the original request. If speaking again restarts the search, even natural dialogue could create more work. Comparing similar requests with and without spoken updates should examine repeated commands, abandonment during waits and audio cost per completed request. Those observations can establish whether saying more actually helps.
The cover reuses this site’s Gemini 3.8 Flash / Flash Cyber launch artwork; it is not a product image of the new Live models.
Sources:
- Google: Gemini 3.8 Live and Extended Thinking announcement
- Google: Gemini API release notes and general availability
- Google: Background thinking and interaction status in the Live API
- Google: Extended Thinking model and migration requirements
- Google: Live API capabilities and preview status
- Google: Gemini API pricing
- Google DeepMind: Gemini 3.8 Live evaluation methodology
- SiliconANGLE: Google’s new Gemini 3.8 Live speech models
Related Articles
Google Launches Gemini 3.7 Flash: Half-Price Until Year-End, With Agent Costs Still Tied to Retry Rates
Google positions Gemini 3.7 Flash as a workhorse for coding and AI agents, with higher vendor benchmarks and temporary half-price access, while architecture, training methods, and production retry rates remain undisclosed.
Google Unveils Gemini 4 Argon: Code Modernization Needs the Right Baseline
Gemini 4 Argon raises the output limit to 1M tokens, initially for trusted cyber defenders. Its libgav1 example shows iterative improvement of a Rust port, but the 2.7-fold speedup is not measured against the optimized C++ original.