Transcription
Models that turn spoken audio into text.
The transcription models behind audio transcription. Each is priced per minute of audio and documents the languages it covers and whether it separates speakers and marks timestamps. Name one directly when a job needs a particular language, diarization, or a lower cost.
10 models
DDeepgram Nova-3Transcription
Deepgram
$0.0058/ min of audio
February 2025
Released
RealtimeLow latency
OGPT TranscribeTranscription
OpenAI
$10.00/ 1M tokens
March 2025
Released
XGrok STTTranscription
X.ai
$0.10/ hour
2026
Released
Low cost
QQwen3 ASR 1.7BTranscription
Qwen· viaDeepInfra
$0.00045/ min
2026
Released
Open sourceMultilingual
EScribe v1Transcription
ElevenLabs
$0.0067/ min
July 2025
Released
EScribe v2Transcription
ElevenLabs
$0.0067/ min
2026
Released
MVoxtral Mini 3BTranscription
Mistral· viaDeepInfra
$0.0010/ min
July 2025
Released
Open source
OWhisper Large v3Transcription
OpenAI· viaDeepInfra
$0.00045/ min
November 2023
Released
Open sourceLow cost
OWhisper Large v3 TurboTranscription
OpenAI· viaDeepInfra
$0.00020/ min
October 2024
Released
Open sourceFastLow cost
OWhisper-1Transcription
OpenAI
$0.006/ minute
September 2022
Released
Open sourceLow cost