AI-FirstResults-DrivenDigital & AI Agency 9800 Richmond Ave, Houston, TX 77042 Start Your Brief

Glossary · AI & Agents

Deepgram

Definition: Deepgram is a voice AI company whose API provides fast, accurate speech-to-text transcription for both real-time streaming and pre-recorded audio. Its Nova model family supports dozens of languages, features like speaker diarization and redaction, and low-latency streaming, making it a common backend for call analytics, captioning, and conversational voice agents.

Official source: Deepgram

Overview

What Deepgram is

Deepgram is a voice AI platform best known for its speech-to-text (automatic speech recognition) API. It transcribes both live streaming audio and pre-recorded files, and also offers text-to-speech and audio understanding capabilities for building voice applications.

Its speech recognition is delivered through the Nova model family; Nova-3, released in early 2025, added real-time multilingual transcription and self-serve keyterm prompting. Deepgram is developer-focused, billing largely by audio minutes, and is used to add voice to products without training models in-house.

How Deepgram works

Applications send audio to Deepgram's API, either as a live stream over a WebSocket for real-time results or as a file for batch processing, and receive transcribed text with word-level timestamps. Options include speaker diarization, punctuation and formatting, keyterm prompting for domain vocabulary, and redaction of sensitive entities.

Deepgram emphasizes low latency and throughput, which suits real-time use cases like live captioning and voice agents. It provides SDKs across common languages and integrates with voice-agent frameworks, so its transcription can feed a language model whose response is then spoken back to the caller.

Why Deepgram matters

Fast, accurate transcription is the foundation of most voice applications, from call-center analytics to real-time voice assistants. Deepgram is often chosen for its speed, streaming performance, and per-minute pricing, and it offers specialized options such as a medical model for domain accuracy.

It is one of several speech providers; OpenAI, Google, Amazon, and AssemblyAI offer competing speech-to-text services with different accuracy, language coverage, latency, and cost profiles. The right choice depends on your languages, real-time needs, and budget, and it is worth benchmarking on your own audio.

FAQ

Deepgram — common questions

What is Deepgram used for?

Deepgram is used for speech-to-text transcription in call analytics, meeting and media captioning, voice assistants, and conversational voice agents. It handles both real-time streaming audio and batch processing of pre-recorded files.

What is Deepgram Nova-3?

Nova-3 is Deepgram's speech-to-text model family released in early 2025. It added real-time multilingual transcription, self-serve keyterm prompting for custom vocabulary, and accuracy improvements over earlier Nova models.

How is Deepgram priced?

Deepgram bills primarily by audio processed, typically per minute, with different rates for streaming versus batch and for specific models. It offers pay-as-you-go usage with free starter credits; check current rates, as pricing can change.

Keep exploring

Related terms

Need Deepgram done right?

Book a free consultation and we’ll map the fastest, most cost-effective path for your project.

Book a free consultation