Google DeepMind has detailed Gemini 3.8 Audio, a new set of models built specifically for real-time dialogue where speed and natural interaction matter.
The family includes Gemini 3.8 Audio Live and a Live Extended Thinking version, combining native multimodal processing with the Gemini 3 generation’s reasoning capabilities.
Why build a separate real-time audio model?
Voice assistants have a different performance problem from text chat. A model can be extremely intelligent, but long pauses make spoken interaction feel unnatural. Google says Gemini 3.8 Audio is optimized for high-volume, latency-sensitive tasks such as real-time conversation.
Extended thinking gives developers another option
The extended-thinking variant is intended for interactions that need more reasoning. That creates a tradeoff developers can choose depending on whether immediate response speed or deeper processing matters more for a particular task.
Multimodal reasoning is part of the design
Google describes the models as natively multimodal. That matters because future assistants are increasingly expected to combine speech with other information rather than treating audio as a separate transcription step.
Known limitations still matter
Google’s model card documents limitations, safety mitigations and evaluation information. As with any voice AI, developers should consider transcription errors, accents, noisy environments and the possibility of incorrect model responses before using it for consequential tasks.
Where this is heading
Real-time audio is becoming a major competitive area for AI platforms. Faster dialogue can make AI useful in tutoring, accessibility, customer service and hands-free assistants, but natural speech should not be confused with guaranteed accuracy.
Source: Google DeepMind Gemini 3.8 Audio model card, published September 15, 2026.