Scribe v2
ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Plain-English overview
Scribe v2 turns long audio or video into structured transcripts with language detection, word timestamps, speaker diarization, and audio tags for events beyond speech. Keyterm prompting helps teams preserve names, products, and vocabulary that general recognition models often miss.
ElevenLabs offers file transcription and a separately priced realtime route. The broader platform also provides voice and agent products, which can simplify one-vendor audio pipelines, although each capability has its own billing and data controls.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Scribe v2
ElevenLabs
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route.
Global and data-residency options depend on the ElevenLabs plan, endpoint, and enterprise configuration.
Check live availabilityData and training
ElevenLabs lets eligible API customers manage data-use preferences. Zero-data-retention through enable_logging=false is limited to qualifying enterprise configurations and supported models, including Scribe v2; verify the live plan and regional endpoint.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists. Scribe v2 turns long audio or video into structured transcripts with language detection, word timestamps, speaker diarization, and audio tags for events beyond speech. Keyterm prompting helps teams preserve names, products, and vocabulary that general recognition models often miss.
Scribe v2 was released on January 9, 2026 according to the cited provider materials.
Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route. The access routes listed in this guide are ElevenLabs and ElevenLabs API.
$0.22 / audio hour for Scribe v2. ElevenLabs lists separate pricing for batch Scribe v2 and Scribe v2 Realtime. Entity detection and redaction can be billed separately.
Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route. Global and data-residency options depend on the ElevenLabs plan, endpoint, and enterprise configuration.
ElevenLabs lets eligible API customers manage data-use preferences. Zero-data-retention through enable_logging=false is limited to qualifying enterprise configurations and supported models, including Scribe v2; verify the live plan and regional endpoint. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Open full comparisonMicrosoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Open full comparisonGoogle Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Open full comparisonAssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Open full comparison