Glossary · AI & Agents
Deepgram
Overview
What Deepgram is
Deepgram is a voice AI platform best known for its speech-to-text (automatic speech recognition) API. It transcribes both live streaming audio and pre-recorded files, and also offers text-to-speech and audio understanding capabilities for building voice applications.
Its speech recognition is delivered through the Nova model family; Nova-3, released in early 2025, added real-time multilingual transcription and self-serve keyterm prompting. Deepgram is developer-focused, billing largely by audio minutes, and is used to add voice to products without training models in-house.
How Deepgram works
Applications send audio to Deepgram's API, either as a live stream over a WebSocket for real-time results or as a file for batch processing, and receive transcribed text with word-level timestamps. Options include speaker diarization, punctuation and formatting, keyterm prompting for domain vocabulary, and redaction of sensitive entities.
Deepgram emphasizes low latency and throughput, which suits real-time use cases like live captioning and voice agents. It provides SDKs across common languages and integrates with voice-agent frameworks, so its transcription can feed a language model whose response is then spoken back to the caller.
Why Deepgram matters
Fast, accurate transcription is the foundation of most voice applications, from call-center analytics to real-time voice assistants. Deepgram is often chosen for its speed, streaming performance, and per-minute pricing, and it offers specialized options such as a medical model for domain accuracy.
It is one of several speech providers; OpenAI, Google, Amazon, and AssemblyAI offer competing speech-to-text services with different accuracy, language coverage, latency, and cost profiles. The right choice depends on your languages, real-time needs, and budget, and it is worth benchmarking on your own audio.
Where we use it
Related Zen in Tech services
How our team puts Deepgram to work in real projects.
FAQ
Deepgram — common questions
What is Deepgram used for?
Deepgram is used for speech-to-text transcription in call analytics, meeting and media captioning, voice assistants, and conversational voice agents. It handles both real-time streaming audio and batch processing of pre-recorded files.
What is Deepgram Nova-3?
Nova-3 is Deepgram's speech-to-text model family released in early 2025. It added real-time multilingual transcription, self-serve keyterm prompting for custom vocabulary, and accuracy improvements over earlier Nova models.
How is Deepgram priced?
Deepgram bills primarily by audio processed, typically per minute, with different rates for streaming versus batch and for specific models. It offers pay-as-you-go usage with free starter credits; check current rates, as pricing can change.
Need Deepgram done right?
Book a free consultation and we’ll map the fastest, most cost-effective path for your project.
Knowledge hub
From our knowledge hub
All articles →How to Reduce AI Voice Agent Latency
How to cut AI voice agent latency to sub-second, human-like turn-taking: where lag comes from (STT, LLM, TTS, network) and the fixes that actually work.
Read · 7 min →AI AutomationAutomating Lead Follow-Up and Onboarding for Coaches and Agencies
Follow-up automation for coaches and agencies: respond to leads in minutes, qualify prospects before calls, and automate onboarding so you focus on clients.
Read · 6 min →AI AutomationAI Automation for Enrollment Inquiries: Answer Every Family Fast
Slow replies lose enrollments. See how AI chatbots and automated follow-up answer every inquiry fast, day or night, and hand warm leads to your team.
Read · 6 min →