Muse Voice Transcribe
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Plain-English overview
Muse Voice Transcribe is designed for speech that arrives continuously rather than only as a finished upload. It can process small audio chunks, decide when an utterance has ended, label many speakers, and use language, keyword, and context hints to improve names or specialist vocabulary.
Meta advertises broad multilingual training while identifying a smaller set of specifically verified languages. That distinction is useful for procurement: test the exact accents, code-switching patterns, noise, and speaker overlap in your own calls before treating the wider training set as production coverage.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
Muse Voice Transcribe
Meta
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through Meta Model API, Meta AI for Mac, and Muse Code during the Model API public preview.
Meta does not enumerate one model-specific country or processing-region list on the public launch page; use the current developer-account terms.
Check live availabilityData and training
Meta's launch page describes privacy for its public demo, not a complete Model API retention or training-use commitment. Review current Model API terms and enterprise controls before sending sensitive audio.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription. Muse Voice Transcribe is designed for speech that arrives continuously rather than only as a finished upload. It can process small audio chunks, decide when an utterance has ended, label many speakers, and use language, keyword, and context hints to improve names or specialist vocabulary.
Muse Voice Transcribe was released on September 1, 2026 according to the cited provider materials.
Available through Meta Model API, Meta AI for Mac, and Muse Code during the Model API public preview. The access routes listed in this guide are Meta Model API and Meta.
$0.18 / audio hour. Meta lists this Model API rate on the reviewed catalog. Product allowances in Meta AI for Mac or Muse Code can differ.
Available through Meta Model API, Meta AI for Mac, and Muse Code during the Model API public preview. Meta does not enumerate one model-specific country or processing-region list on the public launch page; use the current developer-account terms.
Meta's launch page describes privacy for its public demo, not a complete Model API retention or training-use commitment. Review current Model API terms and enterprise controls before sending sensitive audio. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Open full comparisonGoogle Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Open full comparisonAssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Open full comparisonElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Open full comparison