Dubbing AI wanted Voice Agent to feel like a person, not a voice assistant
Dubbing AI had already built a platform around voice. Its users could change their voices in real time, clone them, refine accents, and translate between languages. Voice Agent extended that product in a new direction: the AI could now speak on the user's behalf.
The use cases were immediate. A gamer could have the agent respond during a conversation. Someone who was too busy to make a phone call could have the agent do it for them. Someone who was temporarily unavailable could still communicate in real time.
But the experience only worked if the voice sounded real.
A noticeable pause would make the interaction feel artificial. A flat or robotic response would do the same. For Dubbing AI, real-time TTS wasn't simply a matter of generating speech quickly. The voice needed to sound natural, carry emotion, preserve the speaker's identity, and work across the languages its global users needed.
The fastest model wasn't enough if it sounded flat
Dubbing AI evaluated multiple TTS providers against the demands of the Voice Agent experience. The team narrowed the decision to five criteria: naturalness, emotional depth, voice cloning quality, latency, and multilingual support.
The evaluation created a difficult tradeoff. Models that performed well on latency could lose emotional expression. Models that sounded expressive could introduce delays that disrupted conversation. Others performed well in one language but fell short when users switched languages.
For a product serving more than 10 million users across gaming, entertainment, and commercial use cases, being excellent at only one part of the experience wasn't enough.
Dubbing AI needed a model that could clear every bar at once.
Fish Audio delivers exceptional voice naturalness, rich emotional expression, and reliable low-latency TTS that perfectly underpin our core Voice Agent product experience.

Tiange Ling
CEO, Dubbing AI
Fish Audio was the only model to meet all five requirements
Fish Audio stood out because it delivered the combination Dubbing AI needed for Voice Agent: natural speech, emotional depth, strong voice cloning, low latency, and multilingual support.
The distinction mattered because the criteria were interconnected. Naturalness made the voice believable. Emotional expression made it feel human. Voice cloning preserved the identity of the speaker. Low latency kept the conversation moving. Multilingual support allowed the same experience to work across Dubbing AI's global user base.
Fish Audio supported more than 80 languages with native code-switching, while maintaining the responsiveness required for real-time interactions.
Naturalness
Speech that sounds human, not synthesized
Emotional depth
Expression that carries through every nuance
Voice cloning
Cloned voices that preserve the speaker’s identity
Low latency
Real-time responses without noticeable delay
Multilingual support
80+ languages with native code-switching
Fish Audio gives Dubbing AI room to scale Voice Agent
Voice Agent is now moving from a technical possibility to a product Dubbing AI can bring to its global user base. With Fish Audio handling real-time TTS, Dubbing AI can focus on building new ways for its users to communicate while maintaining the speed, expression, and voice quality the experience requires.
The result is a voice agent that doesn't just respond. It sounds like someone you want to keep talking to.