What happens when a voice leaves American English
HeyGen's creator base is international. Australian, Indian, British, South African, Singaporean, Irish, Scottish, Filipino, and Caribbean creators all speak English, and all sound different doing it.
The text-to-speech layer HeyGen had in production produced accurate American English and pulled everything else toward it. Prosody flattened. Vowels neutralized. The accent that made a creator recognizable to their own audience came back sounding close to neutral American. The pattern held across most of the market, and for a reason: models trained American-English-first and adapted afterward tend to perform well on American English and degrade predictably on everything else.
The loudest ticket in the queue
Creators noticed it right away. Accent similarity became a recurring complaint: their voices no longer sounded like themselves. More importantly, the problem was concentrated in the exact audience HeyGen was trying to grow. Global creators needed their voices to travel across languages and accents without losing the identity that made them recognizable in the first place.
HeyGen tested the market on one question
The evaluation was not a feature comparison. HeyGen wasn’t asking which model had the lowest latency, the most expressive delivery, or the longest list of supported languages. They reduced the evaluation to a much simpler question: when a creator cloned their voice, did it still sound like them?
More specifically, HeyGen tested whether the model could preserve the accent that made a voice recognizable in the first place. Could a British creator keep their British accent? Could a voice move between languages without defaulting to an Americanized version of itself?
That became the bar. A voice could be fast, expressive, and technically impressive, but if the accent changed, the voice failed the test.
Fish is a fantastic voice model for maximizing voice similarity for very unique voices.

John Wu
Audio Engineering Manager, HeyGen
Similarity won over latency and cost
HeyGen selected Fish Audio because it was the strongest match against that criterion. Speed and price mattered, but neither was decisive. Once the models cleared the performance bar, the differentiator was simple: which one could deliver the most faithful voice across different accents?
Fish Audio stood out because it preserved the characteristics of the source voice rather than smoothing them toward a common baseline. That gave HeyGen the consistency it needed to expand its voice product to more creators and markets without asking creators to compromise on how they sounded.
The result was more than a better voice model. It removed a barrier to using HeyGen across languages and gave the company a stronger foundation for its global creator strategy.
Full production volume from the first week
HeyGen deployed S2 Pro through the Fish Audio cloud API. Every video a creator publishes on a Fish-powered voice runs through the model. There was no need for a staged rollout or pilot cohort. The integration went to production scale immediately.
The result: voices that stay themselves
Since migrating to Fish Audio, HeyGen has had zero accent-similarity issues in production. In voice tests, creators chose Fish Audio 3:1 over the alternative.
Today, Fish Audio powers voice experiences for HeyGen’s global creator base, helping every voice sound unmistakably like its creator, wherever they speak.