Skip to main content
Microsoft Foundry

Azure-Speech-Speech-to-text

Microsoft
Version: 1

Azure Speech - Speech to text

Transcribes streaming or recorded audio into readable text across 140+ languages and dialects. Accuracy can be further optimized with custom models for your specialized use cases.

Speech to text offers the following core features:

Real-time speech to text: Instant transcription with intermediate results for streaming audio inputs.

Fast transcription: Fastest synchronous file-based processing for situations with predictable latency.

Batch transcription: Efficient processing for large volumes of prerecorded audio files.

LLM speech (preview): Transcribe and translate audio files using LLM-enhanced speech models, with improved quality and support for prompt tuning.

Custom speech: Fine-tune models with enhanced accuracy for specific domains and use cases.

Training cut-off date

This information is not available.

Input formats

Real-time speech to text: 8khz/16-kHz mono audio, PCM, ALAW, MULAW, G722

Fast transcription and Batch transcription: WAV, MP3, OPUS/OGG, FLAC, WMA, AAC, ALAW in WAV container, MULAW in WAV container, AMR, WebM, SPEEX

Supported language

Speech to text supports over 140 locales .

Supported Azure regions

List of supported Azure regions .

Sample JSON response

Please refer to the sample JSON for real-time transcription , fast transcription , or batch transcription according to your usage.

Model architecture

This information is not available.

Quick facts

PublisherMicrosoft
TypeAutomatic speech recognition, Speech to text
LifecycleGenerally available (GA)
Input typeaudio
Output typetext