Google Launches New Voice AI Models for Building Real-Time Conversational Apps

Google is expanding its artificial intelligence capabilities with new voice-focused AI models designed to make real-time conversations between people and applications faster, more natural, and more interactive. Through the Gemini Live API, Google has introduced Gemini 3.8 Live and Gemini 3.5 Transcribe, giving developers new tools for creating voice assistants, transcription platforms, customer-service agents, and other conversational AI applications.

The latest models focus on several important areas of modern AI development, including low-latency voice communication, multilingual conversations, real-time speech recognition, visual understanding, speaker identification, and automatic language detection.

For developers building the next generation of conversational applications, these capabilities could make it easier to create AI systems that can listen, understand context, respond naturally, and interact with users in real time.

What Are Google’s New Gemini Voice AI Models?

Google’s new models are designed for two different but closely connected purposes.

Gemini 3.8 Live is focused primarily on real-time voice conversations. It allows applications to communicate with users through streaming audio while handling API and tool calls in the background.

Gemini 3.5 Transcribe, on the other hand, is designed specifically for speech-to-text applications. It can convert spoken language into written text for both live conversations and previously recorded audio.

Together, these models provide developers with building blocks for creating more advanced real-time conversational AI applications.

Gemini 3.8 Live for Real-Time Voice Conversations

Gemini 3.8 Live is Google’s conversational voice model built for applications where fast responses are important. Instead of requiring an application to wait for an entire interaction to finish before generating a response, the model is designed to support streaming audio conversations.

One of its notable capabilities is the ability to perform API and tool calls in the background while continuing to stream audio responses. This can help applications maintain a smoother conversational experience.

For example, a voice-based customer support assistant could potentially retrieve information from a database or call an external service while continuing the interaction instead of creating long pauses between responses.

This type of architecture can be particularly useful for:

  • AI-powered customer support
  • Virtual assistants
  • Voice-enabled websites
  • Interactive learning applications
  • Healthcare communication systems
  • Automated business support
  • Voice-based productivity tools

Real-Time Visual Understanding

Another important capability of Gemini 3.8 Live is its ability to work with visual information alongside spoken conversations.

According to Google, the model can analyze live images and video at up to one frame per second. This means developers can build applications where the AI understands not only what a user says but also what the user is showing to the system.

For example, a user could point a camera toward an object and ask an AI assistant about it. The application could combine the user’s spoken question with visual information to provide a context-aware response.

This creates opportunities for multimodal AI applications, where voice, images, and video work together rather than being handled as separate inputs.

Support for 97 Languages

Multilingual communication is another major feature of Gemini 3.8 Live.

Google says the model supports 97 languages and is designed to handle realistic accents. It can also automatically switch languages during a conversation.

Automatic language switching could be particularly valuable for international businesses and applications serving multilingual users. Instead of requiring users to manually select a language before starting a conversation, an AI system can potentially identify the language being used and continue the interaction accordingly.

This capability may also help developers build voice assistants for global audiences without creating separate conversational systems for every supported language.

Gemini 3.8 Live Pricing and Session Limits

Google has also published usage pricing for Gemini 3.8 Live. Audio input starts at approximately $0.005 per minute, while audio output starts at around $0.018 per minute.

The model also has session-duration limitations. Audio-only sessions can last for up to 15 minutes, while sessions involving both audio and video can last for up to 2 minutes.

These limits are important considerations for developers planning applications that require long-running voice conversations.

Gemini 3.5 Transcribe: Google’s Speech-to-Text AI Model

While Gemini 3.8 Live focuses on interactive voice conversations, Gemini 3.5 Transcribe is designed specifically for speech recognition and transcription.

Speech-to-text technology is used across many industries, from meeting transcription and customer support to media production and accessibility tools. Google’s new model is designed to provide developers with real-time and recorded-audio transcription capabilities.

Gemini 3.5 Transcribe can be accessed through two API options: the Live API and the Interactions API.

Real-Time Streaming Transcription

The Live API is designed for applications that require speech recognition while a conversation is taking place.

Google says it can provide captions with sub-second latency, making it suitable for scenarios where users need text to appear almost immediately after speech.

Potential applications include:

  • Live captions
  • Online meetings
  • Voice assistants
  • Call-center applications
  • Accessibility services
  • Live events
  • Real-time translation systems

Fast transcription can make voice applications more accessible because users can receive both audio and text during an interaction.

Transcribing Recorded Audio

The Interactions API is intended for pre-recorded audio files of up to one hour.

It also supports features such as speaker diarization and word-level timestamps.

Speaker diarization separates a conversation into sections associated with different speakers. This is particularly useful when transcribing meetings, interviews, podcasts, customer calls, and group discussions.

Word timestamps can provide additional value by identifying when specific words were spoken. Developers can use this information for searchable transcripts, subtitles, video editing, content analysis, and other applications.

Support for 85 Languages

Gemini 3.5 Transcribe supports 85 languages, with automatic language detection and support for regional accents.

Google reports a word error rate of approximately 4% for streaming transcription and around 2.6% for non-streaming transcription.

Word error rate is an important measurement for speech-recognition systems because it indicates how frequently the generated transcript differs from the spoken words. Actual accuracy can vary depending on factors such as background noise, speaker characteristics, audio quality, accents, and the complexity of the conversation.

How Developers Can Access the New Models

Developers can access Google’s new voice AI capabilities through Google AI Studio and the Gemini API.

This gives developers an opportunity to experiment with the models and integrate voice functionality into their own applications.

The availability of dedicated models for conversational audio and transcription could also simplify application architecture. Instead of building every component of a voice application independently, developers can use Google’s AI infrastructure for speech understanding and generation.

AI-Generated Audio and SynthID Watermarking

Google is also addressing the growing issue of identifying AI-generated audio.

According to Google, audio generated by Gemini Live models includes SynthID watermarking. The technology is intended to help identify AI-generated content and provide an additional mechanism for distinguishing synthetic audio from naturally recorded speech.

As voice AI becomes increasingly realistic, technologies that help identify synthetic media could become more important for reducing confusion and addressing misinformation.

What Google’s New Voice AI Models Mean for Developers

The introduction of Gemini 3.8 Live and Gemini 3.5 Transcribe highlights the growing shift toward real-time, multimodal, and multilingual AI applications.

Traditional chatbots have primarily relied on typed text. Newer conversational AI systems are moving toward more natural interactions involving speech, visual information, and live context.

With Gemini 3.8 Live, developers can explore applications that combine voice conversations with visual understanding. Gemini 3.5 Transcribe provides a dedicated solution for converting live and recorded speech into text.

Together, these capabilities could support new products across customer service, education, business communication, accessibility, entertainment, and productivity.

The Future of Real-Time Conversational AI

Voice AI is rapidly becoming an important part of the broader artificial intelligence ecosystem. Users increasingly expect digital assistants to understand natural speech, respond quickly, recognize different languages, and maintain context during conversations.

Google’s latest Gemini voice models are aimed directly at these requirements.

The combination of real-time audio streaming, visual context, multilingual support, automatic language detection, speaker diarization, timestamps, and AI-generated audio watermarking gives developers a broader toolkit for building conversational applications.

As these technologies continue to develop, voice-based AI could become a more common interface for interacting with software and digital services.

For developers and businesses, the key opportunity is not simply adding voice to an existing application. It is creating experiences where conversation, context, and AI-powered actions work together naturally in real time.


Discover more from AiTechtonic - AI & Informative News

Subscribe to get the latest posts sent to your email.