Gemini 3.8 Live Extended Thinking
A separate Gemini voice model that works through complex requests in the background while keeping the conversation going.
A separate Gemini voice model that works through complex requests in the background while keeping the conversation going.
Plain-English overview
Extended Thinking combines live speech with background planning and tool use. It can give spoken progress updates while working through a multi-step request, with low, medium or high thinking levels.
Its input and output limits match Gemini 3.8 Live. Google lists the same token rates for both, but extra reasoning can increase usage. GPT-Live 1 instead delegates deeper work to a separately chosen backend agent.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available through the Gemini Live API and Google AI Studio in supported regions.
Use Google's live available-regions list and verify feature restrictions for the selected route.
Check live availabilityData and training
Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
A separate Gemini voice model that works through complex requests in the background while keeping the conversation going. Extended Thinking combines live speech with background planning and tool use. It can give spoken progress updates while working through a multi-step request, with low, medium or high thinking levels.
Gemini 3.8 Live Extended Thinking was released on September 15, 2026 according to the cited provider materials.
Available through the Gemini Live API and Google AI Studio in supported regions. The access routes listed in this guide are Google.
$0.005/min audio input · $0.018/min audio output. Paid-tier rates: $3 audio input and $12 audio output per million tokens. Text is $0.75 input and $4.50 output; image/video input is $1 per million tokens. Thinking and paid search usage can add cost.
Available through the Gemini Live API and Google AI Studio in supported regions. Use Google's live available-regions list and verify feature restrictions for the selected route.
Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
OpenAI's full-duplex voice frontend for natural conversation that delegates deeper reasoning and actions to a separate backend agent.
Open full comparisonGoogle's real-time voice model for responsive conversations with audio, text and visual context.
Open full comparisonOpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
Open full comparisonGoogle's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Open full comparisonElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Open full comparisonInworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Open full comparisonSpaceXAI's current realtime speech-to-speech model for sub-second conversational agents with tool access.
Open full comparison