Leadership and Companies
Google opens Gemini voice models for real-time agents and transcription
Google released new Gemini models that let developers build voice agents and transcription tools through its API and AI Studio.
What happened
Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking through the Gemini API and Google AI Studio. The speech-to-speech models can maintain conversations while calling tools, processing visual inputs and handling complex reasoning. Gemini 3.5 Transcribe supports 85-plus languages and recorded a 4.0% streaming Word Error Rate. Live API pricing starts at $0.005 per minute for audio input and $0.018 for output.
Why it matters
Developers can build voice agents that execute tasks during conversations, while the pricing and partner integrations support deployment in real-world products.
Source: Google Blog