Edited by Editor-in-Chief, The Indus Pulse 16 Sept 2026, 07:32 PM 3 min readai
Google Unveils Gemini 3.8 Live Models With Asynchronous Tools and Extended Thinking
Google has launched two new voice-focused artificial intelligence models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, engineered to support near real-time conversational reasoning, background task execution, and live visual grounding. Announced as part of an expansion of Google's real-time voice infrastructure, the models allow voice agents to continue talking with users while running APIs and asynchronous tool calls, rather than pausing conversations during processing.
The standard Gemini 3.8 Live model is optimized for scale and cost efficiency, featuring the ability to process visual inputs in near real-time and automatically detect and switch between 97 supported languages mid-conversation while maintaining accent consistency. Alongside the voice models, Google introduced Gemini 3.5 Transcribe, a speech-to-text transcription model supporting more than 85 languages with smart transcription features designed to remove filler words and handle alphanumeric sequences.
Extended Thinking for Complex Workflows
The Gemini 3.8 Live Extended Thinking model is engineered for multi-step workflows, handling deeper intelligence and providing progress updates while maintaining an uninterrupted conversational flow. According to Google, the model utilizes natural verbal cues, such as acknowledging prompts with conversational filler phrases, to manage ongoing tasks in the background as users progress.
The architecture has delivered top-tier benchmark results across industry evaluations. Google reported that the Extended Thinking model topped Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6, scored 97.7 percent on Big Bench Audio, and recorded 68.6 percent on the tau-Voice benchmark alongside 35.1 percent on Sierra's tau-Voice-banking test.
Developer and Enterprise Availability
For software developers and enterprise clients, both new AI models are accessible via the Gemini API and Google AI Studio, as well as in private preview within Gemini Enterprise and upcoming availability for customer experience platforms. Developer ecosystems including Agora, LangChain, LiveKit, and Vercel have integrated the Gemini Live API to handle real-time voice streaming infrastructure.
Pricing for the developer API is set at $0.005 per minute for audio input and $0.018 per minute for audio output. Google is also working with enterprise partners including Salesforce, Genspark, and Lumeris to deploy the voice models across business workflows.
Consumer Rollout and Content Watermarking
Consumer access to the models spans multiple Google platforms. Gemini 3.8 Live has been rolled out for users in Search Live, while the Extended Thinking model is available within Gemini Live. All Google AI users can access the advanced model across Gmail and Keep, with Google AI Pro and Ultra subscribers receiving expanded access in Docs, Gmail, and Keep.
To address safety and authenticity concerns, all audio generated using the Gemini 3.8 models is watermarked using Google's SynthID technology, an inaudible marker designed to identify synthetic audio content and mitigate the dissemination of misinformation.
Shares of parent company Alphabet Class A declined by 1.26 percent to close at $344.98 on September 15 following the announcement, trading at a price-to-earnings multiple of 17.3 with a market capitalization of approximately $4.2 trillion.
Sources & Citations
The Indus Pulse is committed to accuracy and transparency.

