Google released two new voice AI models: Gemini 3.8 Live (optimized for high-volume, cost-efficient deployments) and Gemini 3.8 Live Extended Thinking (for complex multi-step tasks). Both solve the latency problem in voice agents by enabling near-real-time conversation—processing happens in parallel with speaking, tools execute silently in background, and the model provides acknowledgment signals before final answers. Key features include near-real-time multimodal input processing, auto-detection of language switches across 97 languages, and background API execution without conversation gaps. Available now via Gemini API and in limited preview for enterprise customers.
← Back to all articles