Azure-Speech-Speech-to-text
Azure Speech - Speech to text
Transcribes streaming or recorded audio into readable text across 140+ languages and dialects. Accuracy can be further optimized with custom models for your specialized use cases.
Speech to text offers the following core features:
Real-time speech to text: Instant transcription with intermediate results for streaming audio inputs.
Fast transcription: Fastest synchronous file-based processing for situations with predictable latency.
Batch transcription: Efficient processing for large volumes of prerecorded audio files.
LLM speech (preview): Transcribe and translate audio files using LLM-enhanced speech models, with improved quality and support for prompt tuning.
Custom speech: Fine-tune models with enhanced accuracy for specific domains and use cases.
Training cut-off date
This information is not available.
Input formats
Real-time speech to text: 8khz/16-kHz mono audio, PCM, ALAW, MULAW, G722
Fast transcription and Batch transcription: WAV, MP3, OPUS/OGG, FLAC, WMA, AAC, ALAW in WAV container, MULAW in WAV container, AMR, WebM, SPEEX
Supported language
Speech to text supports over 140 locales .
Supported Azure regions
List of supported Azure regions .
Sample JSON response
Please refer to the sample JSON for real-time transcription , fast transcription , or batch transcription according to your usage.
Model architecture
This information is not available.