Model comparison
Chirp 3 vs Scribe v2
Compare Chirp 3 and Scribe v2 using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Model comparison
Compare Chirp 3 and Scribe v2 using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
Quick take
Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Check the Chirp 3 table for your precise language and region. Speaker labels, timestamps, adaptation, and batch behavior have route-specific constraints that a single coverage number can hide.
ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Do not assume batch and realtime are the same route or price. Zero-data-retention is conditional rather than automatic, so verify the plan, endpoint, region, and enable_logging behavior before processing sensitive audio.
Compare the published facts
Values use each provider's own published units and limits. A blank means the provider did not publish a directly comparable value in the sources reviewed.
| Speech recognition & transcription | Chirp 3 | Scribe v2 |
|---|---|---|
| Transcription priceCurrent provider price per audio hour or minute for the listed processing route. | $0.016/min standard · $0.003/min dynamic batch | $0.22 / audio hour; realtime listed separately |
| Live or batchWhether the model handles realtime streams, uploaded recordings, or both. | Streaming, synchronous, and batch | Batch and long audio; separate realtime route |
| LanguagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed. | 85+ languages and locales | 90+ languages |
| Speaker labelsWhether the model identifies who spoke and any published speaker limit. | Supported for a documented subset of languages | Up to 32 speakers |
| TimestampsAvailable word-, segment-, or utterance-level timing information. | Word-level timestamps with route constraints | Word-level timestamps |
| Vocabulary controlKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms. | Speech adaptation, custom vocabulary, language detection, denoising | Up to 1,000 keyterms; audio tags and language detection |
| Where to use itDirect API, cloud catalog, application, or regional endpoint documented by the provider. | Google Cloud Speech-to-Text V2 | ElevenLabs Speech-to-Text API |
How to choose
Start with the job you need to complete, then validate cost, access, and policy details on your exact provider route.
Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.
Provider and API links
ElevenLabs lets eligible API customers manage data-use preferences. Zero-data-retention through enable_logging=false is limited to qualifying enterprise configurations and supported models, including Scribe v2; verify the live plan and regional endpoint.
Provider and API links
Frequently asked questions
Chirp 3: Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions. Scribe v2: ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Consider Chirp 3 when your priority is Multilingual cloud transcription. Consider Scribe v2 when your priority is Multilingual media transcription. Test both with your own data and provider route before committing.
No. This comparison aligns provider-published facts for the Speech recognition & transcription category. It does not claim a universal winner or combine incompatible third-party benchmark scores.