Back to Speech recognition & transcription

Model comparison

Chirp 3 vs Universal-3.5 Pro

Compare Chirp 3 and Universal-3.5 Pro using the same provider-sourced speech recognition & transcription rubric. No mystery score and no invented benchmark ranking.

Facts checked September 4, 2026

Estimate your cost

Set your usage. Your estimate updates as you type.

Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.

How this estimate works

Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.

Estimate your cost
ModelEstimated total (USD)
Chirp 3Google CloudNo reviewed rate
Universal-3.5 ProAssemblyAINo reviewed rate

Quick take

Chirp 3

Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.

Best for

  • Multilingual cloud transcription
  • Teams mixing realtime and batch workloads
  • Google Cloud applications that need regional endpoints

Watch out for

Check the Chirp 3 table for your precise language and region. Speaker labels, timestamps, adaptation, and batch behavior have route-specific constraints that a single coverage number can hide.

Universal-3.5 Pro

AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.

Best for

  • Meetings and calls with language switching
  • Developer products needing a dedicated speech API
  • Workflows that benefit from contextual or keyterm prompting

Watch out for

Async and realtime routes have different prices, and retention depends on endpoint and account controls. Verify optional feature charges, EU routing, training opt-out, BAA status, and zero-retention eligibility.

Compare the published facts

Chirp 3 vs Universal-3.5 Pro

Values use each provider's own published units and limits. A blank means the provider did not publish a directly comparable value in the sources reviewed.

Speech recognition & transcriptionChirp 3Universal-3.5 Pro
Transcription priceCurrent provider price per audio hour or minute for the listed processing route.$0.016/min standard · $0.003/min dynamic batch$0.21/hour async · $0.45/hour streaming
Live or batchWhether the model handles realtime streams, uploaded recordings, or both.Streaming, synchronous, and batchUploaded files and realtime streaming
LanguagesProvider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.85+ languages and locales18 languages at launch with native code-switching
Speaker labelsWhether the model identifies who spoke and any published speaker limit.Supported for a documented subset of languagesYes
TimestampsAvailable word-, segment-, or utterance-level timing information.Word-level timestamps with route constraintsWord-level timestamps
Vocabulary controlKeyword boosting, custom spelling, context, prompting, or other ways to improve domain terms.Speech adaptation, custom vocabulary, language detection, denoisingContextual prompting and keyterm prompting
Where to use itDirect API, cloud catalog, application, or regional endpoint documented by the provider.Google Cloud Speech-to-Text V2AssemblyAI US and EU APIs

How to choose

Compare the job, not the hype.

Start with the job you need to complete, then validate cost, access, and policy details on your exact provider route.

Chirp 3

Google says Speech-to-Text content is not used beyond providing the service unless the customer opts into data logging. Streaming and synchronous content is handled in memory; asynchronous results may be retained temporarily as documented.

Universal-3.5 Pro

AssemblyAI's retention and model-training behavior depends on endpoint, contract, BAA status, EU processing, and opt-out settings. Its docs describe zero-data-retention options for eligible realtime use; verify the exact account configuration.

Frequently asked questions

Chirp 3 vs Universal-3.5 Pro FAQ

What is the main difference between Chirp 3 and Universal-3.5 Pro?

Chirp 3: Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions. Universal-3.5 Pro: AssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.

Should I choose Chirp 3 or Universal-3.5 Pro?

Consider Chirp 3 when your priority is Multilingual cloud transcription. Consider Universal-3.5 Pro when your priority is Meetings and calls with language switching. Test both with your own data and provider route before committing.

Is this Chirp 3 vs Universal-3.5 Pro comparison based on Cody benchmarks?

No. This comparison aligns provider-published facts for the Speech recognition & transcription category. It does not claim a universal winner or combine incompatible third-party benchmark scores.