Offer expires in 47H 59M 59S

We're turning one 🎂

50% Off everything.

Claim Offer

Customer stories

Real-time Voice Agent TTS for 10 million users.

How Dubbing AI built its Voice Agent on Fish Audio — the only TTS that delivered all five capabilities a real-time agent demands: naturalness, emotional depth, voice cloning quality, low latency, and multilingual support.

Industry
Consumer · Gaming · Entertainment
Region
Global
Use case
Voice Agent (real-time TTS)
Audience
10M+ users
Deployment
Cloud API · desktop & mobile
Status
Voice Agent in beta
10M+

Across gaming, streaming, and entertainment — the scale at which a voice agent has to feel real in real time, in the user's language.

Dubbing AI

Tiange Ling

CEO of Dubbing AI

"Fish Audio delivers exceptional voice naturalness, rich emotional expression, and reliable low-latency TTS that perfectly underpin our core Voice Agent product experience."

CEO of Fish Audio

Rissa Cao

"Voice agents are the use case where every voice AI tradeoff gets exposed. You can't be fast and flat, or expressive and slow. The agent has to sound real, in real time, in the user's language. Dubbing AI is building the version of this that's hardest: a voice that speaks for the user when they can't or don't want to speak themselves. The bar is identity-level realism, and that's the bar we built S2 Pro to clear."
Dubbing AI

About Dubbing AI — 10 million users across
voice creation, cloning, and changing.

Dubbing AI is a local-end AI voice technology product offering one-stop voice creation, voice cloning, and real-time voice changing across desktop and mobile. The platform serves over 10 million users globally across three core audiences: content creators and streamers, individual entertainment users, and commercial users including advertising agencies and media companies.

Dubbing AI is the voice layer for gamers, streamers, and brand creators who need to shape, change, or clone voice in real time across the platforms they live on.

Introducing the Voice Agent
— an agent that speaks for
you.

Dubbing AI's newest feature — the Voice Agent — extends the platform's voice toolkit from voice changing to voice speaking. Instead of changing the user's own voice, the Voice Agent speaks for the user.

The use cases are immediate and human. When a gamer is in the middle of an argument and doesn't want to escalate, the Voice Agent can speak for them. When someone is too tired to make a phone call reservation, the Voice Agent makes it. When someone is shy, occupied with work, or temporarily unavailable to communicate, the Voice Agent enables them to interact with others in real time.

It's an extension of Dubbing AI's product mission across every feature on the platform — voice changing, accent refinement, real-time translation, and now Voice Agent: help people communicate more smoothly and express themselves more effectively. Voice Agent stretches that mission into the situations where speaking for oneself isn't possible or isn't ideal.

For the Voice Agent to work, the AI's voice has to feel real. The hearing party — the gaming opponent, the restaurant host, the person on the other end of the line — should feel like they're talking to a real person, not an obviously synthesized voice. That's where Fish Audio came in.

The challenge of real-time TTS for voice agents: latency vs naturalness.

Voice agents make voice AI's hardest tradeoff visible. Real-time TTS for voice agents has to balance latency against naturalness and emotion, and most providers force a choice. Low-latency models tend to sound flat and machine-like; expressive models tend to introduce processing delays that break conversational flow.

For the Dubbing AI Voice Agent, both factors are equally critical. A Voice Agent that pauses noticeably between user input and spoken response breaks the illusion that the hearing party is talking to a real person. A Voice Agent that responds instantly but sounds robotic breaks the same illusion in a different direction. The deciding factor isn't either alone — it's the combination.

Why Dubbing AI evaluated the TTS market for voice agent infrastructure.

Dubbing AI evaluated multiple TTS audio workflows before settling on Fish Audio. The evaluation criteria mapped directly to the structural demands of TTS for voice agents: naturalness, emotional depth, voice cloning quality, low latency, and multilingual support — five capabilities most providers deliver on two or three of, but rarely all five.

For a Voice Agent serving 10 million users across gaming, entertainment, and commercial use cases, a model that nailed naturalness but failed on multilingual was disqualified. A model that nailed latency but flattened emotion was disqualified. The Voice Agent use case forced an all-five-or-nothing evaluation.

Why Fish Audio won the Voice Agent evaluation — all five criteria.

Fish Audio stood out for the combination Dubbing AI couldn't find anywhere else: all five capabilities at the level a real-time voice agent demands. The single-criterion winners every other provider produced disqualified them from the Voice Agent use case. Fish was the only model that earned its place across every evaluation dimension.

· Naturalness
Voice output that sounds like a real person speaking — not a synthesizer reading.
· Emotional depth
Emotional register that carries across an utterance — the layer most low-latency models flatten.
· Voice cloning quality
Cloned voices that maintain identity across content, important across Dubbing AI's creator and entertainment audiences.
· Low latency
Real-time response without noticeable processing delay — the table-stakes constraint for any conversational agent.
· Multilingual support
80+ languages with native code-switching, required for a Voice Agent serving a global user base.

How Dubbing AI uses Fish Audio for real-time voice agent TTS.

Dubbing AI deploys Fish Audio via the cloud API for real-time text-to-speech generation inside the Voice Agent feature. As users compose the text they want the Voice Agent to speak, Fish converts it to natural, emotionally expressive voice output in real time — across the languages and accents Dubbing AI's global user base requires.

The Voice Agent runs cross-platform on both desktop and mobile, matching the rest of the Dubbing AI platform's coverage. The Voice Agent is preparing for beta release to the platform's gamer audiences first — the user segment with the strongest demand for the use cases Voice Agent was built for. Internal testing results have been very positive heading into beta.

Outcomes from the integration.

Products used: Fish Audio S2 Pro · Text-to-Speech (cloud API)

10M+ users on Dubbing AI's broader platform across gaming, streaming, and commercial creators.

5 of 5 evaluation criteria met by Fish Audio: naturalness, emotional depth, cloning quality, low latency, multilingual.

Voice Agent beta launching to gamer audiences first, with positive internal testing results.

Cross-platform deployment on desktop and mobile, matching the full Dubbing AI surface.

What's next for Dubbing AI and Fish Audio.

As the Voice Agent moves from beta to general availability across Dubbing AI's 10 million users, Fish Audio remains the real-time TTS layer powering the experience. Future expansions of the Voice Agent (into more languages, more situations, and more cross-platform contexts) ship alongside Fish's continuing model improvements.

Fish Audio

Building a voice agent

Talk to our team about real-time TTS that balances naturalness, emotional depth, latency, and multilingual support — the five-way combination voice agents demand.