All models

OpenAI

GPT-Realtime-2.1

OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.

Plain-English overview

What GPT-Realtime-2.1 actually is

GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.

Audio, text, and image input are billed in different token classes. A realistic budget therefore needs captured usage from full conversations, not a simple conversion from audio minutes alone.

Good fit for

  • Tool-using customer or employee voice agents
  • Multimodal realtime assistants
  • OpenAI-native agent stacks

Category comparison

The facts that matter for voice models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Price basisToken-, character-, or minute-based price for the listed access route.
Audio $32 in · $64 out / 1M tokens
Interaction modeSpeech-to-speech, audio-to-audio, or streamed text-to-speech behavior.
Speech-to-speech with text/image context
Published latencyProvider-published measurement with its stated exclusions; not a Cody test.
Not published
LanguagesProvider-documented language coverage or a clear unpublished marker.
Multilingual; exact list not published here
Tools & controlNative function calling, reasoning, or speech-control features.
Function calling and configurable reasoning
Context or input limitDocumented token or character limit where it is meaningful.
128K input · 32K max output

API and provider access

Where to get GPT-Realtime-2.1

Availability

Regions and access stage

Available through OpenAI's Realtime API in supported countries and territories.

Direct access follows OpenAI's live supported-country list.

Check live availability

Data and training

The route matters.

OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

GPT-Realtime-2.1 FAQ

What is GPT-Realtime-2.1?

OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context. GPT-Realtime-2.1 is designed for live conversations where the model must listen, speak, follow instructions, and call tools in one realtime session. OpenAI highlights improved interruption, silence, noise, and alphanumeric handling over the prior version.

When was GPT-Realtime-2.1 released?

The provider does not publish a clear release date for GPT-Realtime-2.1 in the source material reviewed by Cody.

Where can I access GPT-Realtime-2.1?

Available through OpenAI's Realtime API in supported countries and territories. The access routes listed in this guide are OpenAI.

How much does GPT-Realtime-2.1 cost?

$32 audio input · $64 audio output / 1M tokens. Text is $4 input and $24 output per million tokens; image input is $5 per million tokens. Cached-input rates differ.

Where is GPT-Realtime-2.1 available?

Available through OpenAI's Realtime API in supported countries and territories. Direct access follows OpenAI's live supported-country list.

Is my GPT-Realtime-2.1 API data used for training?

OpenAI says API content is not used for training by default. Default abuse-monitoring logs may be retained up to 30 days, with additional controls available to qualifying organizations. The policy belongs to the provider route and account terms, so verify it again before production use.