Explanatory journalism with depth and rigorPTENES
Explosão SolarContext. Not just headlines.Search

Google Launches New Gemini Models for Real-time Voice

New Audio AI Versions Promise Smarter, Multitasking Conversations

Daniele Morais
September 16, 2026 · 2 min read
ShareWhatsAppXFacebook

Last Tuesday, Google announced the availability of three new audio artificial intelligence models: Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.5 Transcribe. The announcement followed the company's blog post stating that these features are already integrated into the Gemini API, Google AI Studio, and products like Gmail, Docs, and Keep.

What are the Gemini 3.8 Live and Extended Thinking Models?

Gemini 3.8 Live brings a performance shift in speech-to-speech capabilities, allowing the voice agent to perform tasks while maintaining the flow of conversation. The Extended Thinking variant adds multi-step reasoning, with “configurable thinking” that processes complex requests in the background and narrates progress to the user.

Practical Applications for Developers and Businesses

Both models are offered via the Live API, priced at US$0.005 per minute for audio input and US$0.018 per minute for output. Google indicates that the solution can be scaled through integration partners such as Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents, who handle the streaming infrastructure.

Gemini 3.5 Transcribe: Accurate Transcription in Over 85 Languages

The Gemini 3.5 Transcribe model, launched last month, focuses on low-latency audio transcription. It achieves a word error rate (WER) of 4.0% in streaming and 2.6% in non-streaming mode, supporting over 85 languages. Features include structured timestamps, speaker diarization, and the ability to transcribe files up to one hour long.

Integration with Google Workspace Products

Beyond serving external developers, the new models power voice features in Google's own applications. Gmail Live enables conversational search in emails, Docs Live offers text generation and editing by voice, and Keep Live transforms spoken ideas into structured lists and notes.

Impact for Brazilian Users

With the expansion of real-time audio understanding capabilities, users in Brazil will be able to interact with more fluid voice assistants, including in Portuguese, leveraging the 97 languages supported by Gemini 3.8 Live. The promise of “conversing with AI more intuitively and intelligently” is reflected in the reduction of friction when performing tasks such as scheduling appointments, searching for information, or generating documents without needing to type.

With information from blog.google, PYMNTS.com, 9to5Google.

Source: blog.google, PYMNTS.com, 9to5Google

#Google#Gemini#Voice AI#API#Workspace#Transcription
Also inPortuguêsEspañol
ShareWhatsAppXFacebook