MAI-Transcribe-2
Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts.
Plain-English overview
MAI-Transcribe-2 is aimed at turning large audio files into usable text quickly. Microsoft pairs multilingual recognition with diarization, word-level timestamps, automatic language detection, and controls for whether the transcript preserves filler words or returns cleaner prose.
Access runs through Microsoft Foundry, which makes the model a natural fit for organizations already managing identity, regions, and data controls in Azure. The launch price is explicitly limited-time, so cost projections should retain a margin for a future standard rate.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
Pricing & comparisons
Set your usage. Your estimate updates as you type.
Uses the recording length in hours: 30 minutes = 0.5 hours. Extra features and minimum charges may change the bill.
MAI-Transcribe-2
Microsoft
Estimated total (USD)
For the usage above · USD · API pricing, not a subscription
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
API and provider access
Availability
Available through Microsoft Foundry for batch and long-form speech recognition.
Deployment availability follows the live Microsoft Foundry model catalog and the Azure region chosen by the customer.
Check live availabilityData and training
Microsoft says prompts, outputs, embeddings, and training data submitted to Foundry Models are not available to model providers and are not used to train foundation models without permission. Retention and abuse-monitoring details depend on the deployed service.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Microsoft's fast batch transcription model for long recordings, speaker labels, word timing, keyword biasing, and clean or verbatim transcripts. MAI-Transcribe-2 is aimed at turning large audio files into usable text quickly. Microsoft pairs multilingual recognition with diarization, word-level timestamps, automatic language detection, and controls for whether the transcript preserves filler words or returns cleaner prose.
MAI-Transcribe-2 was released on September 3, 2026 according to the cited provider materials.
Available through Microsoft Foundry for batch and long-form speech recognition. The access routes listed in this guide are Microsoft AI and Microsoft Foundry.
$0.10 / audio hour for the limited-time preview rate. Microsoft labels this as limited-time pricing. Confirm the live Foundry catalog before forecasting sustained production cost.
Available through Microsoft Foundry for batch and long-form speech recognition. Deployment availability follows the live Microsoft Foundry model catalog and the Azure region chosen by the customer.
Microsoft says prompts, outputs, embeddings, and training data submitted to Foundry Models are not available to model providers and are not used to train foundation models without permission. Retention and abuse-monitoring details depend on the deployed service. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Meta's streaming speech-recognition model for live captions, long audio, many speakers, code-switching, and domain-aware transcription.
Open full comparisonGoogle Cloud's general speech-to-text model for streaming, file, and low-cost dynamic-batch transcription across many languages and regions.
Open full comparisonAssemblyAI's accuracy-focused transcription model for files and live audio, with code-switching, speaker labels, timestamps, and contextual prompting.
Open full comparisonElevenLabs' multilingual speech-to-text model for detailed transcripts with many speakers, word timing, audio events, and large custom keyterm lists.
Open full comparison