Inworld AI
Inworld Realtime TTS-2
Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Inworld AI
Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Plain-English overview
Inworld Realtime TTS-2 turns text into streamed speech while using earlier audio turns to keep a character's delivery consistent. Developers can describe the desired mood or delivery in natural language, switch languages, and use voice cloning where their rights and account setup permit it.
Pricing is based on characters rather than audio tokens, which makes scripts easier to estimate. Inworld reports sub-200ms median time to first audio, but network distance, application work, and the upstream language model still affect the complete response time.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API.
Standard region coverage is not listed on the model page; enterprise plans advertise EU and India data-residency options.
Check live availabilityData and training
Inworld advertises zero data retention as an enterprise add-on, but its currently published ZDR support page names only TTS-1.5 Mini and Max. Confirm TTS-2 coverage directly before sending regulated or confidential text.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Inworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages. Inworld Realtime TTS-2 turns text into streamed speech while using earlier audio turns to keep a character's delivery consistent. Developers can describe the desired mood or delivery in natural language, switch languages, and use voice cloning where their rights and account setup permit it.
Inworld Realtime TTS-2 was released on August 31, 2026 according to the cited provider materials.
Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API. The access routes listed in this guide are Inworld AI, Inworld AI, Inworld AI.
$25 / 1M characters on demand. Published subscription rates fall to $20, $17.50, $15, and $12.50 per million characters by tier, with enterprise pricing advertised as low as $5.
Generally available through Inworld's REST TTS API and OpenAI-compatible Realtime API. Standard region coverage is not listed on the model page; enterprise plans advertise EU and India data-residency options.
Inworld advertises zero data retention as an enterprise add-on, but its currently published ZDR support page names only TTS-1.5 Mini and Max. Confirm TTS-2 coverage directly before sending regulated or confidential text. The policy belongs to the provider route and account terms, so verify it again before production use.
Research trail
Keep comparing
OpenAI · Stable
OpenAI's realtime speech-to-speech reasoning model for tool-using voice agents that also need text and image context.
Read guideGoogle · Preview
Google's preview audio-to-audio model for low-latency dialogue with multimodal awareness, thinking, search grounding, and function calling.
Read guideElevenLabs · Stable
ElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Read guide