Voice cloning for AI video, 3-to-1 over alternatives.
How HeyGen migrated from ElevenLabs to Fish Audio to solve voice similarity issues with non-American English accents — and watched users pick Fish three-to-one in side-by-side comparisons.
How HeyGen migrated from ElevenLabs to Fish Audio to solve voice similarity issues with non-American English accents — and watched users pick Fish three-to-one in side-by-side comparisons.
Selection rate for Fish Audio over alternative TTS providers in side-by-side user comparisons on HeyGen.
John Wu
Audio Engineering Manager for HeyGen
"Fish is a fantastic voice model for maximizing voice similarity for very unique voices."
CEO of Fish Audio
Rissa Cao

"HeyGen is where voice cloning quality matters most. Knowledge creators don't need a voice that sounds like them: they need a voice that sounds like them, accent and all. That bar, the bar that flattens most TTS models when they leave American English, is exactly what we built S2 Pro to clear. Watching HeyGen's users pick Fish three-to-one in side-by-side tests is the validation we built the model for."

HeyGen is the leading AI video platform for knowledge-based creators — the educators, course builders, business communicators, and content creators who turn expertise into video at scale. The platform lets creators produce professional AI-generated video without studio equipment, professional voice talent, or weeks of production time.
Voice sits at the center of every video HeyGen produces. For a knowledge creator whose audience already knows their voice, the AI voice has to be their voice — not a generic narrator, not a close approximation, but the actual voice with the actual accent that audiences recognize. That bar is the bar voice cloning for AI video has to clear.
Before evaluating Fish Audio, HeyGen was using ElevenLabs for the TTS layer of its AI video generation pipeline.
The platform worked. The voices were intelligible. But one structural problem kept surfacing in user
feedback: voice similarity issues, especially for creators with non-American English accents.
Knowledge creators speak with the accents they have. Australian creators sound Australian. Indian creators sound
Indian. British, South African, Singaporean, Irish, Scottish, Filipino, Caribbean — the accent landscape of global
English isn't an edge case for HeyGen, it's the shape of the international creator base.
A TTS that produces strong American English but flattens every other accent into something closer to neutral
American breaks the core promise of voice cloning for AI video: that the AI video sounds like you. HeyGen's users
with non-American English accents were complaining about accent maintenance. The voice clones weren't
preserving the accent identity creators expected.


HeyGen evaluated the entire TTS market to solve the non-
American English accent problem. TTS for non-American English accents wasn't a feature comparison — it was an existential evaluation criterion for the platform's international creator base.
Across providers tested, the gap between American English voice cloning quality and non-American English voice cloning quality was consistent — and consistent in the wrong direction. Most TTS models were trained American-English-first, then adapted for other accents. The adaptation showed up as flattened prosody, neutralized vowels,
and a distinct loss of accent identity in cloned voices.
Fish Audio earned HeyGen's selection on one criterion above all others: voice cloning quality for unique voices, particularly voices with non-American English accents. The Fish S2 Pro model preserves the accent identity of
the source voice rather than normalizing toward an American English baseline — the specific failure mode
HeyGen's testing had surfaced across other providers.
For HeyGen, this isn't a technical detail. It's the difference between a knowledge creator's audience hearing the
creator they recognize and hearing a synthesized stranger. Voice cloning quality, not latency or cost, was the
deciding factor.

HeyGen deploys Fish Audio via the cloud API, generating TTS for the videos HeyGen's customers produce on the platform. The integration is production-
scale: every video that a knowledge creator publishes through HeyGen using a Fish-powered voice runs through Fish's S2 Pro model.
The most telling user signal isn't praise — it's silence on what was previously the loudest complaint. Three signals from HeyGen's creator base since the migration:
Products used: Fish Audio S2 Pro · Text-to-Speech (cloud API)
3:1 user selection rate for Fish Audio over alternative TTS providers in side-by-side comparisons on HeyGen.
Non-American English accent complaints ceased across HeyGen's international creator base.
Entire TTS market evaluated before selecting Fish Audio — accent preservation was the deciding factor.
Migrated from ElevenLabs to Fish Audio for voice cloning quality.
Talk to our team about voice similarity for unique voices, non-American English accent preservation, and production-scale cloud API integration.