Chirp 3
Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Plain-English overview
Chirp 3 sits inside Google Cloud Speech-to-Text V2, so it supports several operational patterns rather than one fixed app. Teams can use it for live streams, short synchronous requests, or asynchronous batches and can add language detection, speech adaptation, and denoising where the selected region and language support them.
Its feature matrix is not uniform across every language. Diarization and some adaptation options apply to documented subsets, and endpoint location affects availability. The right comparison is therefore an exact language, region, and request mode—not the headline language count alone.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Chirp 3
Google Cloud
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes.
Features and supported languages vary by Google Cloud region; choose a regional endpoint from the live Chirp 3 documentation.
Check live availabilityData and training
Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions. Chirp 3 sits inside Google Cloud Speech-to-Text V2, so it supports several operational patterns rather than one fixed app. Teams can use it for live streams, short synchronous requests, or asynchronous batches and can add language detection, speech adaptation, and denoising where the selected region and language support them.
Chirp 3 was released on October 13, 2025 according to the cited provider materials.
Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes. The access routes listed in this guide are Google Cloud.
$0.016/min standard · $0.003/min dynamic batch. These are current Google Cloud Speech-to-Text rates for the relevant recognition routes; eligible volume tiers and cloud contracts can change effective cost.
Generally available in Google Cloud Speech-to-Text V2 with streaming, synchronous, and batch recognition routes. Features and supported languages vary by Google Cloud region; choose a regional endpoint from the live Chirp 3 documentation.
Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Open full comparisonMicrosoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Open full comparisonAssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Open full comparisonElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Open full comparison