Resources

Top AI Models for Speech Recognition & Transcription

Speech-to-text models for meetings, calls, captions, media, and voice products. Compare hourly cost, live or batch processing, languages, speaker labels, timestamps, vocabulary controls, and API access.

Showing 1–5 of 5 models

Transcription

Google Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.

Transcription price
$0.016/min standard · $0.003/min dynamic batch
Live or batch
Streaming, synchronous, and batch
View model guide
Transcription

ElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.

Transcription price
$0.22 / audio hour; realtime listed separately
Live or batch
Batch and long audio; separate realtime route
View model guide

What to compare

Transcription price
Current provider price per audio hour or minute for the listed processing route.
Live or batch
Whether the model handles realtime streams, uploaded recordings, or both.
Languages
Provider-published language coverage, separating trained or advertised coverage from specifically verified languages where needed.

How to use this library

Compare the job, not the hype.

A curated cross-category set of frontier, specialist, and category-defining models with practical API information for real product decisions.

Cody does not run a universal quality leaderboard. Provider claims are labeled, pricing excludes taxes and optional tools, and production buyers should verify the live provider terms before choosing a model.

  1. 01

    Use first-party documentation for technical limits, lifecycle, pricing, regional access, and data policies whenever it exists.

  2. 02

    Compare models only inside the same category and retain each provider's measurement basis.

  3. 03

    Show Not published when a provider has not published a comparable fact instead of estimating it.

  4. 04

    Treat prices and availability as dated snapshots and link every page to live provider documentation.

Speed keeps its units

Voice latency, text throughput, video render time, and world-model frame rate are not interchangeable.

Unknown stays unknown

A missing release date, country list, or latency number is labeled as unpublished instead of estimated.

Every fact has a date

Model pages link to the underlying source and show when pricing, access, and policies were last checked.

Put the model to work

Pick the model, then give it a better brief.

Use Cody's expanded prompt playbooks and curated agent skills to turn a model choice into a repeatable workflow.