GPT-Realtime-2.1
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
Plain-English overview
GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.
Audio, text, and image input are billed in different token classes. A realistic budget therefore needs captured usage from full conversations, not a simple conversion from audio minutes alone.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available through OpenAI's Realtime API in supported countries and territories.
Direct access follows OpenAI's live supported-country list.
Check live availabilityData and training
OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context. GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.
The provider does not publish a clear release date for GPT-Realtime-2.1 in the source material reviewed by Cody.
Available through OpenAI's Realtime API in supported countries and territories. The access routes listed in this guide are OpenAI.
$32 audio input · $64 audio output / 1M tokens. Text is $4 input and $24 output per million tokens; image input is $5 per million tokens. Cached-input rates differ.
Available through OpenAI's Realtime API in supported countries and territories. Direct access follows OpenAI's live supported-country list.
OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
OpenAI's full-duplex voice frontend for natural conversation that delegates deeper reasoning and actions to a separate backend agent.
Open full comparisonGoogle's real-time voice model for responsive conversations with audio, text and visual context.
Open full comparisonA separate Gemini voice model that works through complex requests in the background while keeping the conversation going.
Open full comparisonGoogle's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Open full comparisonElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Open full comparisonInworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Open full comparisonSpaceXAI's current realtime speech-to-speech model for sub-second conversational agents with tool access.
Open full comparison