All models
Inworld AI

Inworld Realtime TTS-2

Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.

Plain-English overview

What Inworld Realtime TTS-2 actually is

Inworld Realtime TTS-2 turns text into streamed speech while using earlier audio turns to keep a character's delivery consistent. Developers can describe the desired mood or delivery in natural language, switch languages, and use voice cloning where their rights and account setup permit it.

Pricing is based on characters rather than audio tokens, which makes scripts easier to estimate. Inworld reports sub-200ms median time to first audio, but network distance, application work, and the upstream language model still affect the complete response time.

Good fit for

  • Realtime companions and support voices
  • Multilingual experiences with one consistent voice
  • Creative voice direction and authorized cloning

Category comparison

The facts that matter for voice models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Price basis
$25 / 1M chars on demand; volume discountsToken-, character-, or minute-based price for the listed access route.
Interaction mode
Steerable realtime TTS with prior-audio contextSpeech-to-speech, audio-to-audio, or streamed text-to-speech behavior.
Published latency
Under 200ms median TTFAProvider-published measurement with its stated exclusions; not a Cody test.
Languages
100+ with mid-utterance switchingProvider-documented language coverage or a clear unpublished marker.
Tools & control
Voice direction, design, cloning, non-verbalsNative function calling, reasoning, or speech-control features.
Context or input limit
Not publishedDocumented token or character limit where it is meaningful.

API and provider access

Where to get Inworld Realtime TTS-2

Availability

Regions and access stage

Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API.

Standard region coverage is not listed on the model page; enterprise plans advertise EU and India data-residency options.

Check live availability

Data and training

The route matters.

Inworld advertises zero data retention as an enterprise add-on, but its currently published ZDR support page names only TTS-1.5 Mini and Max. Confirm TTS-2 coverage directly before sending regulated or confidential text.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

Inworld Realtime TTS-2 FAQ

What is Inworld Realtime TTS-2?

Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages. Inworld Realtime TTS-2 turns text into streamed speech while using earlier audio turns to keep a character's delivery consistent. Developers can describe the desired mood or delivery in natural language, switch languages, and use voice cloning where their rights and account setup permit it.

When was Inworld Realtime TTS-2 released?

Inworld Realtime TTS-2 was released on August 31, 2026 according to the cited provider materials.

Where can I access Inworld Realtime TTS-2?

Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API. The access routes listed in this guide are Inworld AI.

How much does Inworld Realtime TTS-2 cost?

$25 / 1M characters on demand. Published subscription rates fall to $20, $17.50, $15, and $12.50 per million characters by tier, with enterprise pricing advertised as low as $5.

Where is Inworld Realtime TTS-2 available?

Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API. Standard region coverage is not listed on the model page; enterprise plans advertise EU and India data-residency options.

Is my Inworld Realtime TTS-2 API data used for training?

Inworld advertises zero data retention as an enterprise add-on, but its currently published ZDR support page names only TTS-1.5 Mini and Max. Confirm TTS-2 coverage directly before sending regulated or confidential text. The policy belongs to the provider route and account terms, so verify it again before production use.

Related comparisons

Head-to-head comparisons

  1. Inworld Realtime TTS-2 vs Gemini 3.8 Flash TTS

    Google's speech-generation model for expressive narration, character voices and dialogue in 130 languages.

    Open full comparison
  2. Inworld Realtime TTS-2 vs Gemini 3.8 Flash-Lite TTS

    Google's lower-cost speech-generation model for read-aloud features and high-volume audio in 101 languages.

    Open full comparison
  3. Inworld Realtime TTS-2 vs GPT-Live 1

    OpenAI's full-duplex voice frontend for natural conversation that delegates deeper reasoning and actions to a separate backend agent.

    Open full comparison
  4. Inworld Realtime TTS-2 vs Gemini 3.8 Live

    Google's real-time voice model for responsive conversations with audio, text and visual context.

    Open full comparison
  5. Inworld Realtime TTS-2 vs Gemini 3.8 Live Extended Thinking

    A separate Gemini voice model that works through complex requests in the background while keeping the conversation going.

    Open full comparison
  6. Inworld Realtime TTS-2 vs GPT-Realtime-2.1

    OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.

    Open full comparison
  7. Inworld Realtime TTS-2 vs Gemini 3.1 Flash Live Preview

    Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.

    Open full comparison
  8. Inworld Realtime TTS-2 vs Eleven v3 Conversational

    ElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.

    Open full comparison
  9. Inworld Realtime TTS-2 vs Grok Voice Think Fast 2.0

    SpaceXAI's current realtime speech-to-speech model for sub-second conversational agents with tool access.

    Open full comparison