Gemini 3.1 Flash Live Preview
Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Plain-English overview
Gemini 3.1 Flash Live accepts live audio alongside text, images, and video, then returns text and audio. It is aimed at conversations that benefit from acoustic nuance and visual context rather than simple text-to-speech playback.
Google publishes both token rates and per-minute audio equivalents. This makes early budgeting easier, but search grounding and multimodal inputs can still add separate usage.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Preview access through the Gemini Live API and Google AI Studio in supported regions.
Use Google's live Gemini API region list and confirm preview eligibility.
Check live availabilityData and training
Google says paid Gemini API content is not used to improve its products; unpaid service data generally can be. Confirm billing status and current terms before sending sensitive conversations.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling. Gemini 3.1 Flash Live accepts live audio alongside text, images, and video, then returns text and audio. It is aimed at conversations that benefit from acoustic nuance and visual context rather than simple text-to-speech playback.
Gemini 3.1 Flash Live Preview was released on March 1, 2026 according to the cited provider materials.
Preview access through the Gemini Live API and Google AI Studio in supported regions. The access routes listed in this guide are Google.
$0.005/min audio input · $0.018/min audio output. Google's paid-tier equivalents; text input/output are $0.75/$4.50 per million tokens and image/video input is $0.002 per minute equivalent.
Preview access through the Gemini Live API and Google AI Studio in supported regions. Use Google's live Gemini API region list and confirm preview eligibility.
Google says paid Gemini API content is not used to improve its products; unpaid service data generally can be. Confirm billing status and current terms before sending sensitive conversations. The policy belongs to the provider route and account terms, so verify it again before production use.
Research trail
Keep comparing
OpenAI · Stable
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
Read guideElevenLabs · Stable
ElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Read guideInworld AI · Stable
Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Read guide