HeyGen

HeyGen creators choose Fish Audio three-to-one for voice cloning

HeyGen is the AI video platform for knowledge creators, used by educators, course builders, and business communicators to produce professional video without a studio or voice talent. The company moved its voice layer to Fish Audio to preserve the accents of its international creator base.

Products used
Fish Audio S2 ProText-to-Speech APIVoice Cloning
Migrated from
ElevenLabs

What happens when a voice leaves American English

HeyGen's creator base is international. Australian, Indian, British, South African, Singaporean, Irish, Scottish, Filipino, and Caribbean creators all speak English, and all sound different doing it.

The text-to-speech layer HeyGen had in production produced accurate American English and pulled everything else toward it. Prosody flattened. Vowels neutralized. The accent that made a creator recognizable to their own audience came back sounding close to neutral American. The pattern held across most of the market, and for a reason: models trained American-English-first and adapted afterward tend to perform well on American English and degrade predictably on everything else.

The loudest ticket in the queue

Creators noticed it right away. Accent similarity became a recurring complaint: their voices no longer sounded like themselves. More importantly, the problem was concentrated in the exact audience HeyGen was trying to grow. Global creators needed their voices to travel across languages and accents without losing the identity that made them recognizable in the first place.

HeyGen tested the market on one question

The evaluation was not a feature comparison. HeyGen wasn’t asking which model had the lowest latency, the most expressive delivery, or the longest list of supported languages. They reduced the evaluation to a much simpler question: when a creator cloned their voice, did it still sound like them?

More specifically, HeyGen tested whether the model could preserve the accent that made a voice recognizable in the first place. Could a British creator keep their British accent? Could a voice move between languages without defaulting to an Americanized version of itself?

That became the bar. A voice could be fast, expressive, and technically impressive, but if the accent changed, the voice failed the test.

Fish is a fantastic voice model for maximizing voice similarity for very unique voices.

John Wu

Audio Engineering Manager, HeyGen

Similarity won over latency and cost

HeyGen selected Fish Audio because it was the strongest match against that criterion. Speed and price mattered, but neither was decisive. Once the models cleared the performance bar, the differentiator was simple: which one could deliver the most faithful voice across different accents?

Fish Audio stood out because it preserved the characteristics of the source voice rather than smoothing them toward a common baseline. That gave HeyGen the consistency it needed to expand its voice product to more creators and markets without asking creators to compromise on how they sounded.

The result was more than a better voice model. It removed a barrier to using HeyGen across languages and gave the company a stronger foundation for its global creator strategy.

Full production volume from the first week

HeyGen deployed S2 Pro through the Fish Audio cloud API. Every video a creator publishes on a Fish-powered voice runs through the model. There was no need for a staged rollout or pilot cohort. The integration went to production scale immediately.

The result: voices that stay themselves

Since migrating to Fish Audio, HeyGen has had zero accent-similarity issues in production. In voice tests, creators chose Fish Audio 3:1 over the alternative.

Today, Fish Audio powers voice experiences for HeyGen’s global creator base, helping every voice sound unmistakably like its creator, wherever they speak.

Read more customer stories

Fish Audio

Ready when you are

Talk to our team about your deployment. We'll come prepared.