If you've used voice AI in the past year, you've probably noticed something: conversations feel different now. The awkward pauses are gone. The AI responds almost instantly, like talking to a human.
This isn't just faster hardware. It's a fundamental shift in how voice AI works. Native audio LLMs have changed everything.
The Problem with Traditional Voice AI
For years, voice AI followed the same pattern:
[Your Voice] → Speech-to-Text → LLM → Text-to-Speech → [AI Voice]
Each step takes time:
- Speech-to-Text (STT): 150-250ms to transcribe audio
- LLM Processing: 200-400ms to understand and generate response
- Text-to-Speech (TTS): 150-250ms to synthesize voice
Total: 500-900ms minimum — often over a second with network latency.
But worse than latency was what got lost. When you convert speech to text, you lose:
- Tone: Is "fine" agreeable or sarcastic?
- Emotion: Is the speaker frustrated, confused, excited?
- Emphasis: Which words matter most?
- Hesitation: Is there uncertainty behind the words?
The AI was essentially reading a transcript, not hearing a conversation.
Enter Native Audio-to-Audio
Native audio LLMs process audio directly. No intermediate text step:
[Your Voice] → Native Audio LLM → [AI Voice]
The model takes in raw audio waveforms and outputs raw audio waveforms. It "hears" and "speaks" natively, the way humans do.
How It Works (Technical Overview)
Traditional LLMs are trained on text tokens. Native audio LLMs are trained on audio tokens—discrete representations of sound rather than words.
Training process:
- Audio is segmented into short frames (typically 10-20ms)
- Each frame is converted to a learned representation (audio token)
- The model learns relationships between audio token sequences
- Output is generated as audio tokens, then reconstructed to audio
The key insight: Just as text LLMs predict "the next word," audio LLMs predict "the next sound." But unlike text, sounds carry emotional information, emphasis, and timing.
The Result
| Metric | Traditional Pipeline | Native Audio |
|---|---|---|
| Latency | 700-1000ms | 377-500ms |
| Emotional understanding | None | Yes |
| Tone adaptation | None | Yes |
| Natural pacing | No | Yes |
| Code-switching (Hindi-English) | Poor | Excellent |
Available Native Audio Models
Gemini Live 2.5 HD (Google)
Google's most advanced native audio model, released late 2025.
Key specifications:
- 30 HD studio-quality voices
- 24 natively supported languages
- Affective dialog (emotional detection and response)
- 377ms latency with Vertex AI backend
- Function calling support
Standout feature: Affective dialog. The model detects caller emotions (frustrated, confused, happy, anxious) and adjusts its response tone automatically. A frustrated caller gets empathy; a happy caller gets matched enthusiasm.
Best for: Customer service, healthcare, collections, any conversation where emotion matters.
Gemini Live 3.1 (Google)
The latest iteration with improved reasoning capabilities.
Key specifications:
- Same 30 HD voices as 2.5
- 131K token context window
- Better multi-turn reasoning
- Improved instruction following
- Slightly higher latency than 2.5
Standout feature: Extended context window. Can reference much earlier parts of long conversations without losing track.
Best for: Complex multi-turn conversations, technical support, detailed explanations.
OpenAI Realtime (OpenAI)
OpenAI's native audio implementation, bringing GPT-4o quality to voice.
Key specifications:
- 8 premium voices (Alloy, Ash, Ballad, Coral, Echo, Sage, Shimmer, Verse)
- GPT-4o reasoning capabilities
- Robust function calling
- ~500ms latency
- Realtime Mini variant for cost savings
Standout feature: GPT-4o-level reasoning in voice form. For conversations requiring complex logic, multi-step reasoning, or technical explanations, the reasoning quality is unmatched.
Best for: Technical support, complex B2B sales, anything requiring sophisticated reasoning.
OpenAI Realtime Mini (OpenAI)
Cost-optimized version of Realtime.
Key specifications:
- Same 8 voices
- Reduced reasoning capability vs full Realtime
- ~75% lower cost
- Same latency profile
Standout feature: Native audio at the lowest cost. When you need the latency benefits but don't need maximum reasoning.
Best for: High-volume outbound campaigns, simple qualification calls, notification calls.
When to Use Native Audio vs Traditional Pipeline
Native audio isn't always the best choice. Here's when to use each:
Use Native Audio When:
-
Latency is critical
- Sales calls where quick responses build rapport
- Support calls where pauses feel dismissive
- Any conversation where natural pacing matters
-
Emotional context matters
- Collections (de-escalation)
- Healthcare (empathy)
- Complaints (acknowledgment)
-
Code-switching is common
- Hindi-English (Hinglish) conversations
- Spanish-English (Spanglish) in US
- Any bilingual market
-
Voice naturalness is priority
- Brand-representing calls
- High-value customer interactions
Use Traditional Pipeline When:
-
You need specific voices
- Custom cloned voices
- Specific regional accents not in native models
- Character voices for entertainment
-
Language isn't natively supported
- Some languages have better dedicated TTS
- Rare languages with limited native model training
-
Maximum transcription accuracy needed
- Legal/compliance recordings
- Medical documentation
- When every word must be precisely captured
-
Cost is the primary driver
- Some traditional setups are cheaper for simple use cases
- Though Realtime Mini is competitive
Implementation Guide
Configuring Native Audio on Edesy
Switching to native audio is a configuration change, not a code change.
Agent settings:
{
"llmProvider": "gemini-live-2.5",
"voice": "Aoede",
"language": "en-US",
"fallbackLLM": "gpt-4o"
}
Available provider values:
gemini-live-2.5- Gemini Live 2.5 HDgemini-live-3.1- Gemini Live 3.1openai-realtime- OpenAI Realtime (full)openai-realtime-mini- OpenAI Realtime Minigemini-live-2.0- Original Gemini Live
Optimizing for Lowest Latency
Gemini Live achieves 377ms on Vertex AI vs 1,578ms on standard API. Edesy automatically routes to Vertex AI when available.
Latency breakdown:
- Audio capture: ~20ms
- Network to Vertex: ~50ms
- Model processing: ~250ms
- Network return: ~50ms
- Audio playback: ~7ms
- Total: ~377ms
Handling Fallback
Configure automatic fallback to traditional pipeline:
{
"llmProvider": "gemini-live-2.5",
"fallbackEnabled": true,
"fallbackLLM": "gpt-4o",
"fallbackSTT": "deepgram-nova",
"fallbackTTS": "google-neural"
}
If native audio fails, calls continue with traditional pipeline. Callers don't notice the switch.
Common Misconceptions
"Native audio is just faster TTS"
No. Native audio eliminates TTS entirely. The model generates audio directly from understanding, not from text. It's a fundamentally different architecture.
"You can't customize native audio voices"
Partially true. You can't clone voices into native audio models (yet). But you choose from 30+ high-quality voices with Gemini Live. For most use cases, this is sufficient.
"Native audio can't do function calling"
False. Both Gemini Live and OpenAI Realtime fully support function calling. The AI can look up order status, book appointments, or integrate with CRMs—all while using native audio.
"Traditional pipeline gives better transcripts"
Sometimes true. If you need word-perfect transcripts for compliance, dedicated STT may be more accurate. But for most conversational use cases, native audio transcription is excellent.
The Future of Voice AI
Native audio is just the beginning. What's coming:
- Personalized voice adaptation - Models that adjust to individual caller preferences
- Multi-speaker handling - Native support for conference calls
- Real-time translation - Native cross-language conversations
- Emotion-guided responses - Beyond detection to true emotional intelligence
The shift from text-mediated to native audio is as significant as the shift from rule-based to neural AI. We're at the start of a new era in voice technology.
Getting Started
Ready to try native audio?
- Sign up for an Edesy account
- Create an agent and select "Gemini Live 2.5 HD" as the LLM provider
- Choose a voice from the 30 available options
- Test with the built-in phone simulator
- Deploy and experience the difference
The 377ms response time isn't a number on a spec sheet—it's something you feel immediately in conversation.
Experience native audio voice AI firsthand. Start a free trial on Edesy and test Gemini Live 2.5 HD or OpenAI Realtime with your use case.