Serviços
Contate-nos

Speech-to-Text Benchmark: Deepgram vs. Whisper

Cem Dilmegani
Cem Dilmegani
atualizado em 21 ago. 2026

We benchmarked the leading speech-to-text (STT) providers, focusing specifically on healthcare applications. Our benchmark used real-world examples to assess transcription accuracy in medical contexts, where precision is crucial.

Speech-to-text benchmark results

Based on both word error rate (WER) and character error rate (CER) results, GPT-4o-transcribe demonstrates the highest transcription accuracy among all evaluated speech-to-text systems. Deepgram Nova-v3 and Gladia also perform strongly, maintaining low error rates across both metrics.

Loading Chart

Methodology

Dataset

We wanted to evaluate the models’ performances in both small and various samples and a long sample, so we conducted two tasks:

Task 1: Healthcare voice data

  • Total number of samples: 100
  • Total duration: 9 minutes and 25 seconds
  • Average duration per sample: 5.65 seconds
  • Content: Healthcare voice data, including medical terminology, patient interactions, and clinical discussions
  • Variety: Different speakers, varying audio quality, and diverse medical contexts spoken in English

Audio specifications:

  • Format: WAV
  • Channels: 1 (Mono)
  • Sample width: 16-bit
  • Sample rate: 16 kHz
  • Consistent bitrate: 256 kbps
  • Duration range: ~4.5 to 11.5 seconds per file

Task 2: An anatomy lecture

  • Total number of samples: 1
  • Total duration: 8 minutes and 35 seconds
  • Content: An anatomy lecture given by a doctor, including medical terminology
  • Variety: One speaker speaks in English in the first half of the video; music plays in the background.

Audio specifications:

  • Format: WAV
  • Channels: 2 (Stereo)
  • Sample width: 16-bit
  • Sample rate: 48 kHz
  • Consistent bitrate: 1536 kbps

Evaluation metrics

We used word error rate (WER) and character error rate (CER) as evaluation metrics for transcription accuracy. Word error rate is calculated as:

WER = (S + D + I) / N

Where:

  • S = Number of substitutions
  • D = Number of deletions
  • I = Number of insertions
  • N = Total number of words in the ground truth

The formula calculates the minimum number of word-level operations needed to transform the hypothesis into the reference, divided by the number of words in the reference. Lower WER indicates better accuracy, with 0% being a perfect match.

The character error rate (CER) is calculated by dividing the total number of character-level errors (including insertions, deletions, and substitutions) by the total number of characters in the reference text.

We used speech-to-text APIs to transcribe audio files to text.

The maximum file size input at one time by the providers is shown in the table:

*Since Vosk runs locally, there is no limit on the input file size. However, long audio files may exceed the beam limit, causing some probabilities to be lost. Therefore, it is recommended to split the files into 1–2 minute segments.

Google MedASR also operates locally and does not impose a maximum file size limit. For optimal performance and resource management, processing long files in smaller segments is recommended.

Note: For providers with smaller file-size limits (such as Google and OpenAI), larger audio files must be split into smaller chunks before processing. We performed that in Task 2.

Speech recognition

Speech recognition enables computers to transcribe audio files into text using machine learning algorithms. A transcription service’s API can be used with various programming languages for batch transcription. These platforms support both real-time and asynchronous transcription.

Speech recognition technology has numerous applications, including transcription, voice assistants, and language translation.

Benefits of using speech recognition for transcription

  • Fast transcription of audio files
  • Time and effort savings
  • Real-time transcription and translation
  • Accessibility for individuals with disabilities
Deixe nossa equipe automatizar um dos seus processos de negócio com agentes de IA, gratuitamente.
Automatizar um processo

How do speech-to-text IA tools work?

The transcription process includes:

  • Audio data is uploaded or streamed to the speech-to-text tool
  • Usage of machine learning algorithms to analyze the audio data and identify patterns in speech
  • The tool converts the speech to text using a speech-to-text engine
  • The transcribed text is then displayed to the user.

Perguntas frequentes

Transcription of audio and video recordings can be used in:
Voice assistants and virtual assistants
Language translation and interpretation
Speech-to-text (ASR) systems for individuals with disabilities

Their pre-trained models enable automatic speech recognition (ASR) for recorded audio and video files. High-accuracy audio transcriptions include automatic punctuation and topic detection.
An open-source engine or a speech recognition provider from a service your company already works with (i.e., Google Cloud, AWS transcribe) can be chosen as the transcription solution for your company’s needs. Some of them also offer gratuito credits, but we recommend caution regarding data security.

A speech-to-text API can help to transcribe audio files into text. Processing and analysis of audio data:
Audio data is processed using techniques such as noise reduction and echo cancellation
The audio data is then analyzed using machine learning algorithms to identify patterns in speech
The algorithms use acoustic models and language models to recognize spoken words and phrases
Converting speech to text using machine learning algorithms:
Machine learning algorithms are trained on large datasets of audio and text data
The algorithms learn to recognize patterns in speech and convert them into text
The algorithms can be fine-tuned and customized for specific use cases and languages

Não perca os nossos benchmarks e insights baseados em dados. O botão abre o Google; selecionar a AIMultiple confirma que deseja ver a AIMultiple com mais frequência nos resultados de pesquisa do Google.
GoogleAdicionar como fonte preferencial

Further reading

Cite este benchmark

Escolha o formato adequado ao local onde você vai publicar. Colar a versão com link no seu CMS preserva o backlink.

Cem Dilmegani and Şevval Alper (2026) - "Speech-to-Text Benchmark: Deepgram vs. Whisper". Publicado on-line em AIMultiple.com. Acessado em 21 Agosto 2026, em: https://aimultiple.com/speech-to-text [Recurso on-line]

Dilmegani, C., & Alper, Ş. (2026, 21 Agosto). Speech-to-Text Benchmark: Deepgram vs. Whisper. AIMultiple. https://aimultiple.com/speech-to-text

@misc{dilmegani2026,
  author = {Dilmegani, Cem and Alper, Şevval},
  title  = {{Speech-to-Text Benchmark: Deepgram vs. Whisper}},
  year   = {2026},
  month  = aug,
  howpublished    = {\url{https://aimultiple.com/speech-to-text}},
  note   = {AIMultiple. Acessado em 21 Agosto 2026}
}
Baixar todos os dados

Resultados e carimbos de data/hora de 12 pontos de dados. Baixe os dados utilizados neste artigo como um arquivo ZIP contendo um arquivo CSV.

Última atualização: 17 Agosto 2026
Baixar
Cem Dilmegani
Cem Dilmegani
Analista Principal
Cem é o analista principal da AIMultiple desde 2017.

O trabalho de Cem na AIMultiple foi citado por publicações globais líderes, incluindo Business Insider, Forbes, Morning Brew e Washington Post, por empresas globais como Deloitte e HPE, ONGs como o World Economic Forum e organizações supranacionais como a European Commission. [1], [2], [3], [4], [5]

Ao longo de sua carreira, Cem atuou como consultor de tecnologia, comprador de tecnologia e empreendedor de tecnologia. Ele aconselhou empresas sobre suas decisões de tecnologia na McKinsey & Company e na Altman Solon por mais de uma década. Ele também publicou um relatório da McKinsey sobre digitalização.

Ele liderou a estratégia de tecnologia e as compras de uma operadora de telecomunicações, reportando-se ao CEO. Ele também liderou o crescimento comercial da empresa de deep tech Hypatos, que atingiu uma receita recorrente anual de 7 dígitos e uma avaliação de 9 dígitos partindo do zero em 2 anos. O trabalho de Cem na Hypatos foi coberto por publicações de tecnologia líderes como TechCrunch e Business Insider.

Cem fala regularmente em conferências internacionais de tecnologia. Ele se formou como engenheiro da computação pela Bogazici University e possui um MBA pela Columbia Business School.
Ver perfil completo
Pesquisado por
Şevval Alper
Şevval Alper
Pesquisadora de IA
Şevval é pesquisadora de IA na AIMultiple. Ela tem experiência anterior de pesquisa em geração de números pseudoaleatórios usando sistemas caóticos.
Şevval se concentra em ferramentas de codificação de IA, agentes de IA e tecnologias quânticas.
Ver perfil completo

Seja o primeiro a comentar

Seu endereço de e-mail não será publicado. Todos os campos são obrigatórios. Os comentários são deixados em seu idioma original.

0/450