All models
ElevenLabs

Scribe v2

ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.

Plain-English overview

What Scribe v2 actually is

Scribe v2 turns long audio or video into structured transcripts with language detection, word timestamps, speaker diarization, and audio tags for events beyond speech. Keyterm prompting helps teams preserve names, products, and vocabulary that general recognition models often miss.

ElevenLabs offers file transcription and a separately priced realtime route. The broader platform also provides voice and agent products, which can simplify one-vendor audio pipelines, although each capability has its own billing and data controls.

Good fit for

  • Multilingual media transcription
  • Recordings with many speakers or sound events
  • Audio stacks already using ElevenLabs

Category comparison

The facts that matter for transcription models

These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.

Transcription price
$0.22 / audio hour; realtime listed separatelyCurrent provider price per audio hour or minute for the listed processing route.
Live or batch
Batch and long audio; separate realtime routeWhether the model handles realtime streams, uploaded recordings, or both.
Languages
90+ languagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.
Speaker labels
Up to 32 speakersWhether the model identifies who spoke and any published speaker limit.
Timestamps
Word-level timestampsAvailable word-, segment-, or utterance-level timing information.
Vocabulary control
Up to 1,000 keyterms; audio tags and language detectionKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms.
Where to use it
ElevenLabs Speech-to-Text APIDirect API, cloud catalog, application, or regional endpoint documented by the provider.

Pricing & comparisons

Estimate your cost

Set your usage. Your estimate updates as you type.

Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.

Scribe v2

ElevenLabs

Estimated total (USD)

No reviewed rate

For the usage above · USD · API pricing, not a subscription

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

API and provider access

Where to get Scribe v2

Availability

Regions and access stage

Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route.

Global and data-residency options depend on the ElevenLabs plan, endpoint, and enterprise configuration.

Check live availability

Data and training

The route matters.

ElevenLabs lets eligible API customers manage data-use preferences. Zero-data-retention through enable_logging=false is limited to qualifying enterprise configurations and supported models, including Scribe v2; verify the live plan and regional endpoint.

This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.

Read the provider policy

Frequently asked questions

Scribe v2 FAQ

What is Scribe v2?

ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists. Scribe v2 turns long audio or video into structured transcripts with language detection, word timestamps, speaker diarization, and audio tags for events beyond speech. Keyterm prompting helps teams preserve names, products, and vocabulary that general recognition models often miss.

When was Scribe v2 released?

Scribe v2 was released on January 9, 2026 according to the cited provider materials.

Where can I access Scribe v2?

Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route. The access routes listed in this guide are ElevenLabs and ElevenLabs API.

How much does Scribe v2 cost?

$0.22 / audio hour for Scribe v2. ElevenLabs lists separate pricing for batch Scribe v2 and Scribe v2 Realtime. Entity detection and redaction can be billed separately.

Where is Scribe v2 available?

Available through the ElevenLabs speech-to-text API and product interfaces; realtime uses a separate model route. Global and data-residency options depend on the ElevenLabs plan, endpoint, and enterprise configuration.

Is my Scribe v2 API data used for training?

ElevenLabs lets eligible API customers manage data-use preferences. Zero-data-retention through enable_logging=false is limited to qualifying enterprise configurations and supported models, including Scribe v2; verify the live plan and regional endpoint. The policy belongs to the provider route and account terms, so verify it again before production use.