OpenAI
GPT-Realtime-2.1
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
OpenAI
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
Plain-English overview
GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.
Audio, text, and image input are billed in different token classes. A realistic budget therefore needs captured usage from full conversations, not a simple conversion from audio minutes alone.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available through OpenAI's Realtime API in supported countries and territories.
Direct access follows OpenAI's live supported-country list.
Check live availabilityData and training
OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context. GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.
The provider does not publish a clear release date for GPT-Realtime-2.1 in the source material reviewed by Cody.
Available through OpenAI's Realtime API in supported countries and territories. The access routes listed in this guide are OpenAI.
$32 audio input · $64 audio output / 1M tokens. Text is $4 input and $24 output per million tokens; image input is $5 per million tokens. Cached-input rates differ.
Available through OpenAI's Realtime API in supported countries and territories. Direct access follows OpenAI's live supported-country list.
OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations. The policy belongs to the provider route and account terms, so verify it again before production use.
Research trail
Keep comparing
Google · Preview
Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Read guideElevenLabs · Stable
ElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Read guideInworld AI · Stable
Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Read guide