Building a Production-Ready Voice AI Agent with <500ms Latency
How I built a high-performance voice AI agent using Node.js, Deepgram, and OpenAI. A deep dive into system architecture, latency optimization techniques, and handling real-time audio streams.
Goal: Create a human-like voice assistant for restaurant reservations. Result: sub-500ms latency, seamless interruption handling, and robust state management. Stack: Node.js, Deepgram (Nova-2), OpenAI (GPT-4), WebSocket streams.
Building a voice AI that feels “natural” is notoriously difficult. The uncanny valley isn’t just about how the voice sounds—it’s about how it responds. A few hundred milliseconds of delay can turn a fluid conversation into an awkward walkie-talkie exchange.
In this post, I’ll walk you through the architecture of my Voice AI Agent, specifically focusing on how I engineered it for sub-500ms latency and high reliability.
The Latency Challenge
In a typical voice interaction loop, latency accumulates at every step:
- VAD (Voice Activity Detection): Waiting for the user to finish speaking.
- STT (Speech-to-Text): Transcribing audio to text.
- LLM Inference: Generating a response token-by-token.
- TTS (Text-to-Speech): Converting text back to audio.
- Network Overhead: Round-trips between services.
If you chain these sequentially, you easily hit 2-3 seconds of lag. That’s unacceptable for a real-time agent.
Architecture Guidelines
To solve this, I moved from a request-response model to a full-duplex streaming architecture.
1. WebSocket-First Design
Instead of HTTP REST requests, the entire pipeline operates over persistent WebSockets. This minimizes handshake overhead and allows for bidirectional data flow.
// Simplified conceptual flow
ws.on('message', (audioChunk) => {
// 1. Immediately stream audio to Deepgram
deepgramLive.send(audioChunk);
});
deepgramLive.on('transcript', (text) => {
// 2. Stream transcript text to LLM
llmChain.stream(text);
});
2. Optimizing Speech-to-Text (STT)
I used Deepgram’s Nova-2 model for its speed and accuracy. The critical configuration here is the endpointing (VAD) setting.
- Too sensitive: It cuts you off while thinking.
- Too relaxed: It introduces massive delays.
I tuned the VAD to a dynamic window of ~500ms and implemented a “Barge-in” mechanism. When the user starts speaking, the system emits a SpeechStarted event which immediately clears the audio buffer and stops the AI from talking.
“Barge-in” is the most important feature for perceived naturalness. If the user interrupts, the AI must shut up immediately. I achieved this by maintaining a server-side state of is_speaking and flushing the TTS audio queue the moment user audio input is detected.
3. Intelligent Filler Words
Even with optimization, LLM inference takes time. To mask this, I built an Intelligent Filler Word System.
When the STT pipeline detects a completed sentence, but the LLM hasn’t generated the first token, the system plays a context-aware filler sound (e.g., “Hmm,” “Let me check,” “One moment”). This buys the system ~800ms of “thinking time” without the user feeling ignored.
The “Brain”: State Management
Managing a conversation isn’t just about text; it’s about state. The agent needs to know:
- Did I just ask a question?
- Is the user confirming an order?
- Did the call drop?
I implemented a finite state machine (FSM) to handle these transitions. This ensures the AI doesn’t hallucinate an order confirmation when the user was just asking about hours.
Performance Metrics
The results of these optimizations were significant:
| Metric | Standard Pipeline | Optimized Pipeline | Improvement |
|---|---|---|---|
| STT Latency | 600ms+ | less than 250ms | 58% |
| TTFB (Time to First Byte) | 800ms+ | ~300ms | 62% |
| End-to-End Latency | 2.5s+ | ~500ms | 5x Faster |
Conclusion
Building a production-ready Voice AI Agent requires 50% AI and 50% distributed systems engineering. By decoupling the components and streaming everything, we can achieve human-level response times.
The code for this project is part of my portfolio and demonstrates that with the right architecture, we can bridge the gap between “chatbot” and “digital assistant.”
- Deepgram API Documentation Speech-to-text API reference
- OpenAI Realtime API Guidelines for low-latency AI interactions
- Effective Node.js WebSocket Handling Node.js documentation for UDP/Datagram sockets